TEAM-Atreides at SemEval-2022 Task 11: On leveraging data augmentation and ensemble to recognize complex named entities in Bangla

bracu.type.groupResearch Publications
datacite.rightsOpen Access
dc.contributor.authorTasnim, Nazia
dc.contributor.authorShihab, Istiak
dc.contributor.authorSushmit, Asif Shahriyar
dc.contributor.authorBethard, Steven
dc.contributor.authorSadeque, Farig
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-06T05:39:51Z
dc.date.available2026-09-06T05:39:51Z
dc.date.issued2022-01-01
dc.description.abstractBiological and healthcare domains, artistic works, and organization names can all have nested, overlapping, discontinuous entity mentions that may be syntactically or semantically ambiguous in practice. Traditional sequence tagging algorithms are unable to recognize these complex mentions because they violate the assumptions upon which sequence tagging schemes are founded. In this paper, we describe our contribution to SemEval 2022 Task 11 on identifying such complex named entities. We leveraged an ensemble of ELECTRA-based models exclusively pretrained on the Bangla language with ELECTRA-based monolingual models pretrained on English to achieve competitive performance. Besides providing a system description, we also present the outcomes of our experiments on architectural decisions, dataset augmentations and post-competition findings. © 2022 Association for Computational Linguistics.
dc.description.versionPublished
dc.format.extent1524 - 1530
dc.identifier.citationTasnim, N., Shihab, Md. I., Shahriyar Sushmit, A., Bethard, S., & Sadeque, F. (2022). TEAM-Atreides at SemEval-2022 Task 11: On leveraging data augmentation and ensemble to recognize complex Named Entities in Bangla. Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022), 1524–1530. https://doi.org/10.18653/v1/2022.semeval-1.209
dc.identifier.doi10.18653/v1/2022.semeval-1.209
dc.identifier.isbn9781955917803
dc.identifier.other2-s2.0-85137554053
dc.identifier.urihttps://hdl.handle.net/10361/29775
dc.language.isoen_US
dc.publisherAssociation for Computational Linguistics (ACL)
dc.relation.hasversion10.18653/v1/2022.semeval-1.209
dc.relation.ispartofSemeval 2022 16th International Workshop on Semantic Evaluation Proceedings of the Workshop
dc.relation.ispartofseriesSemeval 2022 16th International Workshop on Semantic Evaluation Proceedings of the Workshop
dc.relation.urihttps://aclanthology.org/2022.semeval-1.209/
dc.subjectArchitectural decision
dc.subjectArtistic works
dc.subjectBiological domain
dc.subjectCompetitive performance
dc.subjectData augmentation
dc.subjectData ensemble
dc.subjectHealthcare domains
dc.subjectNamed entities
dc.subjectSystem description
dc.subject.lcshBengali language--Data processing.
dc.subject.lcshNatural language processing (Computer science).
dc.titleTEAM-Atreides at SemEval-2022 Task 11: On leveraging data augmentation and ensemble to recognize complex named entities in Bangla
dc.typeConference Paper
person.affiliation.nameShahjalal University of Science and Technology
person.affiliation.nameShahjalal University of Science and Technology
person.affiliation.nameBengali.Ai
person.affiliation.nameThe University of Arizona
person.affiliation.nameBRAC University
person.identifier.scopus-author-id57522347000
person.identifier.scopus-author-id57221687740
person.identifier.scopus-author-id57202978208
person.identifier.scopus-author-id23092920400
person.identifier.scopus-author-id55843529500

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
TEAM-Atreides at SemEval-2022 Task 11 On leveraging data augmentation and ensemble to recognize complex Named Entities in Bangla.pdf
Size:
216.55 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: