Fine-tuning large language models for regional dialect comprehended question answering in Bangla

Citation

M. J. A. Riad et al., "Fine-Tuning Large Language Models for Regional Dialect Comprehended Question answering in Bangla," 2025 IEEE International Students' Conference on Electrical, Electronics and Computer Science (SCEECS), Bhopal, India, 2025, pp. 1-6, doi: 10.1109/SCEECS64059.2025.10940303.

Abstract

For diverse languages like Bangla, maintaining regional dialects can be a major challenge. The dialect from one region can be difficult to understand for people with dialects of another region and thus, automated system to answer the questions of a particular dialect can be helpful. In this paper, we present a new dataset comprising 12,500 sentences from various regional dialects including Chittagong, Noakhali, Sylhet, Barishal, and Mymensingh, alongside their replies in the same dialect. Afterward, we developed dialect-sensitive chatbots through fine-tuning via Low-Rank Adaptation (LoRA). Our comprehensive evaluation of four leading language models - ChatGPT-4o, Claude 3.5 Sonnet, Mistral-7B, and Gemma-2-9B - reveals significant variations in their ability to process regional Bangla dialects. ChatGPT-4o emerged as the top performer with BLEU scores of 53%, followed by Claude 3.5 Sonnet demonstrating a score of 46%, Gemma-2-9B achieving 42%, and Mistral-7B achieving 40%.

Description

Type

Conference Proceeding