Islam, AshiqulArman, MithilaRahman M.M.2026-10-042026-10-042025-01-01A. Islam, M. Arman and M. M. Rahman, "Multi-Stage Fine-Tuning of T5 for Low-Resource Dialects: An AI-Driven Augmentation Framework," 2025 28th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2025, pp. 4532-4537, doi: 10.1109/ICCIT68739.2025.11490101.97983315786712-s2.0-105041670376https://hdl.handle.net/10361/30378Dialectal diversity poses significant challenges for Bengali Natural Language Processing, where most existing models are trained primarily on Standard Bangla, resulting in poor performance on regional dialects and amplifying systemic bias. To address this gap, this work presents a large-scale dataset containing 63,303 sentences across 12 Bengali dialects, which highlights substantial class imbalance across regions. This work proposes a multi-stage fine-tuning framework leveraging the T5 model, combined with advanced data augmentation techniques back-translation and paraphrasing and class-weighted training to enhance representation of underrepresented dialects. Experiments conducted on the BanglaDial corpus demonstrate that the proposed method achieves state-of-the-art performance, with T5 reaching 92.4% accuracy, 93.0% recall, and 92.1% F1-score outperforming strong baselines including RoBERTa, BERT, and GPT-NeoX. The results confirm that imbalance-aware optimization and synthetic data generation significantly improve model fairness and robustness, making this work a step forward toward inclusive, dialect-aware NLP systems for Bengali.6 Pagesen-USFeedsFilteringFiltersProtocolsSwitchesElectronic componentsMachine learningArtificial intelligenceGenerative pre-trained transformerNatural Language Processing (NLP)Machine learning.Bengali language--Dialects.Artificial intelligence.Multi-stage fine-tuning of T5 for low-resource dialects: An AI-driven augmentation frameworkConference Proceeding10.1109/ICCIT68739.2025.11490101