Multi-stage fine-tuning of T5 for low-resource dialects: An AI-driven augmentation framework

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorIslam, Ashiqul
dc.contributor.authorArman, Mithila
dc.contributor.authorRahman M.M.
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-10-04T06:58:41Z
dc.date.available2026-10-04T06:58:41Z
dc.date.issued2025-01-01
dc.description.abstractDialectal diversity poses significant challenges for Bengali Natural Language Processing, where most existing models are trained primarily on Standard Bangla, resulting in poor performance on regional dialects and amplifying systemic bias. To address this gap, this work presents a large-scale dataset containing 63,303 sentences across 12 Bengali dialects, which highlights substantial class imbalance across regions. This work proposes a multi-stage fine-tuning framework leveraging the T5 model, combined with advanced data augmentation techniques back-translation and paraphrasing and class-weighted training to enhance representation of underrepresented dialects. Experiments conducted on the BanglaDial corpus demonstrate that the proposed method achieves state-of-the-art performance, with T5 reaching 92.4% accuracy, 93.0% recall, and 92.1% F1-score outperforming strong baselines including RoBERTa, BERT, and GPT-NeoX. The results confirm that imbalance-aware optimization and synthetic data generation significantly improve model fairness and robustness, making this work a step forward toward inclusive, dialect-aware NLP systems for Bengali.
dc.description.versionPublished
dc.format.extent6 Pages
dc.identifier.citationA. Islam, M. Arman and M. M. Rahman, "Multi-Stage Fine-Tuning of T5 for Low-Resource Dialects: An AI-Driven Augmentation Framework," 2025 28th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2025, pp. 4532-4537, doi: 10.1109/ICCIT68739.2025.11490101.
dc.identifier.doi10.1109/ICCIT68739.2025.11490101
dc.identifier.issn9798331578671
dc.identifier.other2-s2.0-105041670376
dc.identifier.urihttps://hdl.handle.net/10361/30378
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/ICCIT68739.2025.11490101
dc.relation.ispartof2025 28th International Conference on Computer and Information Technology Iccit 2025
dc.relation.ispartofseries2025 28th International Conference on Computer and Information Technology Iccit 2025
dc.relation.urihttps://ieeexplore.ieee.org/document/11490101
dc.subjectFeeds
dc.subjectFiltering
dc.subjectFilters
dc.subjectProtocols
dc.subjectSwitches
dc.subjectElectronic components
dc.subjectMachine learning
dc.subjectArtificial intelligence
dc.subjectGenerative pre-trained transformer
dc.subjectNatural Language Processing (NLP)
dc.subject.lcshMachine learning.
dc.subject.lcshBengali language--Dialects.
dc.subject.lcshArtificial intelligence.
dc.titleMulti-stage fine-tuning of T5 for low-resource dialects: An AI-driven augmentation framework
dc.typeConference Proceeding
person.affiliation.nameUniversity of Science and Technology Chittagong
person.affiliation.nameBRAC University
person.affiliation.nameUniversity of Science and Technology Chittagong
person.identifier.scopus-author-id60676249900
person.identifier.scopus-author-id58144027900
person.identifier.scopus-author-id58777897300

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Full Text Available at Publisher's Site.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: