Which matters more: Model or language? an empirical study in English-Bangla mental health classification
Loading...
Files
Date
Publisher
Institute of Electrical and Electronics Engineers Inc.
Citation
A. Islam, I. A. Rafi, S. Mondal, S. A. S. Rahman and G. R. Alam, "Which Matters More: Model or Language? An Empirical Study in English-Bangla Mental Health Classification," 2025 5th International Conference on Robotics, Automation, and Artificial Intelligence (RAAI), Singapore, Singapore, 2025, pp. 114-118, doi: 10.1109/RAAI67517.2025.11423350.
Abstract
We present an empirical comparison of classical baselines and pretrained transformer encoders for mental health status classification across English and Bangla social media text. Using two public datasets-an English multi-class corpus and a Bangla binary corpus-we evaluate TF-IDF with Logistic Regression and Random Forest against BERT, RoBERTa, DeBERTa, and BanglaBERT under a matched setup with a stratified eighty twenty split and macro F1 for model selection. In English, RoBERTa achieves 81.6% accuracy with a macro F1 of 78.8, while TF-IDF with Logistic Regression reaches 77.3% accuracy. In Bangla, BanglaBERT attains 88.3% accuracy with a macro F 1 of 88.3, and classical baselines surpass several non-Bangla encoders. Findings highlight the value of language models and the importance of classical machine learning models in classifying mental status across different languages.
LC Subject Headings
Description
Publisher Link
Type
Conference Proceeding