Inceptive transformers: Enhancing contextual representations through multi-scale feature learning across domains and languages

bracu.type.groupResearch Publications
datacite.rightsOpen Access
dc.contributor.authorShahriar, Asif
dc.contributor.authorShahriyar, Rifat
dc.contributor.authorSaifur Rahman M.
dc.date.accessioned2026-09-07T03:10:13Z
dc.date.available2026-09-07T03:10:13Z
dc.date.issued2025-01-01
dc.description.abstractEncoder transformer models compress information from all tokens in a sequence into a single [CLS] token to represent global context. This approach risks diluting fine-grained or hierarchical features, leading to information loss in downstream tasks where local patterns are important. To remedy this, we propose a lightweight architectural enhancement: an inception-style 1-D convolution module that sits on top of the transformer layer and augments token representations with multi-scale local features. This enriched feature space is then processed by a self-attention layer that dynamically weights tokens based on their task relevance. Experiments on five diverse tasks show that our framework consistently improves general-purpose, domain-specific, and multilingual models, outperforming baselines by 1% to 14% while maintaining efficiency. Ablation studies show that multi-scale convolution performs better than any single kernel and that the self-attention layer is critical for performance. © 2025 Association for Computational Linguistics.
dc.description.versionPublished
dc.format.extent25833 - 25848
dc.identifier.citationShahriar, A., Shahriyar, R., & Rahman, M. S. (2025). Inceptive transformers: Enhancing contextual representations through multi-scale feature learning across domains and languages. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 25844–25859. https://doi.org/10.18653/v1/2025.emnlp-main.1312
dc.identifier.doi10.18653/v1/2025.emnlp-main.1312
dc.identifier.isbn9798891763326
dc.identifier.other2-s2.0-105040138526
dc.identifier.urihttps://hdl.handle.net/10361/29786
dc.language.isoen_US
dc.publisherAssociation for Computational Linguistics (ACL)
dc.relation.hasversion10.18653/v1/2025.emnlp-main.1312
dc.relation.ispartofEmnlp 2025 2025 Conference on Empirical Methods in Natural Language Processing Proceedings of the Conference
dc.relation.ispartofseriesEmnlp 2025 2025 Conference on Empirical Methods in Natural Language Processing Proceedings of the Conference
dc.relation.urihttps://aclanthology.org/2025.emnlp-main.1312/
dc.rightsfalse
dc.subjectArchitectural enhancement
dc.subjectDown-stream
dc.subjectFeature learning
dc.subjectFine grained
dc.subjectGlobal context
dc.subjectHierarchical features
dc.subjectInformation loss
dc.subjectLocal patterns
dc.subjectMulti-scale features
dc.subjectTransformer modeling
dc.subject.lcshNatural language processing (Computer science).
dc.subject.lcshMachine learning.
dc.subject.lcshArtificial intelligence--Data processing.
dc.subject.lcshComputational linguistics--Methodology.
dc.subject.lcshComputer network architectures.
dc.titleInceptive transformers: Enhancing contextual representations through multi-scale feature learning across domains and languages
dc.typeConference Paper
person.affiliation.nameBangladesh University of Engineering and Technology
person.affiliation.nameBangladesh University of Engineering and Technology
person.affiliation.nameBangladesh University of Engineering and Technology
person.identifier.scopus-author-id57447028700
person.identifier.scopus-author-id54279365200
person.identifier.scopus-author-id60364263800

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Inceptive Transformers Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and Languages.pdf
Size:
6.67 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: