Benchmarking Bangla named entity recognition: Evaluating dialect robustness across sadhu-cholito

Loading...
Thumbnail Image

Publisher

Institute of Electrical and Electronics Engineers Inc.

Citation

S. C. Nadiya, M. Arman and A. Islam, "Benchmarking Bangla Named Entity Recognition: Evaluating Dialect Robustness Across Sadhu-Cholito," 2025 IEEE 4th International Conference on Robotics, Automation, Artificial-Intelligence and Internet-of-Things (RAAICON), Dhaka, Bangladesh, 2025, pp. 715-720, doi: 10.1109/RAAICON69033.2025.11502037.

Abstract

The most confounding issue underlying any robust named entity recognition (NER) for Bangla lies in the stylistic and dialectal variation between classical Sadhu and modern Cholito registers, and across regional varieties. In this work, present a dialect-aware Bangla NER and fine-tune a RoBERTa encoder with style-sensitive preprocessing with BanglaBlend, a 7.3 k-sentence corpus explicitly labeled for Sadhu and Cholito. To recover informal Cholito forms in this work, use the class-conditional loss that preserves signals relevant for span delimiters and apply a long-horizon schedule to stabilize learning in face of the class imbalance. outperformed powerful encoder and seq2seq baselines to set the new state-of-the-art benchmarks on BanglaBlend with RoBERTa 87.30% accuracy, 87.00% precision, 86.00% recall and 86.50% F 1, with comprehensive comparisons of XLMRoBERTa, ERNIE, BanglaBERT, BART, T5, mBERT, DistilBERT, MarianMT. This work find that, although success is encouraging, error analysis first shows two high-level failure modes, span boundary drift under orthographic variation, and label confusion for long-tail entities which is suggestive of a natural path forward through span-level decoding and coverage expansion. Beyond Bangla, it supplies a practical recipe to provide dialect robustness for low-resource languages-equal mixtures of formal and informal variants-and with careful preprocessing and simple label-faithful sequential augmentation, demonstrates state-of-the-art dialect robustness with encoder-only transformers and no custom architectures.

LC Subject Headings

Description

Type

Conference Proceeding