Shahriar, AsifShahriyar, RifatSaifur Rahman M.2026-09-072026-09-072025-01-01Shahriar, A., Shahriyar, R., & Rahman, M. S. (2025). Inceptive transformers: Enhancing contextual representations through multi-scale feature learning across domains and languages. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 25844–25859. https://doi.org/10.18653/v1/2025.emnlp-main.131297988917633262-s2.0-105040138526https://hdl.handle.net/10361/29786Encoder transformer models compress information from all tokens in a sequence into a single [CLS] token to represent global context. This approach risks diluting fine-grained or hierarchical features, leading to information loss in downstream tasks where local patterns are important. To remedy this, we propose a lightweight architectural enhancement: an inception-style 1-D convolution module that sits on top of the transformer layer and augments token representations with multi-scale local features. This enriched feature space is then processed by a self-attention layer that dynamically weights tokens based on their task relevance. Experiments on five diverse tasks show that our framework consistently improves general-purpose, domain-specific, and multilingual models, outperforming baselines by 1% to 14% while maintaining efficiency. Ablation studies show that multi-scale convolution performs better than any single kernel and that the self-attention layer is critical for performance. © 2025 Association for Computational Linguistics.25833 - 25848en-USfalseArchitectural enhancementDown-streamFeature learningFine grainedGlobal contextHierarchical featuresInformation lossLocal patternsMulti-scale featuresTransformer modelingNatural language processing (Computer science).Machine learning.Artificial intelligence--Data processing.Computational linguistics--Methodology.Computer network architectures.Inceptive transformers: Enhancing contextual representations through multi-scale feature learning across domains and languagesConference Paper10.18653/v1/2025.emnlp-main.1312