Which matters more: Model or language? an empirical study in English-Bangla mental health classification
| bracu.type.group | Research Publications | |
| datacite.rights | Metadata Only | |
| dc.contributor.author | Islam, Apu | |
| dc.contributor.author | Rafi, Ishraque Arefin | |
| dc.contributor.author | Mondal, Sudipta | |
| dc.contributor.author | Rahman, Somaya Al Sadia | |
| dc.contributor.author | Alam, Golam Rabiul | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-08-15T12:53:21Z | |
| dc.date.available | 2026-08-15T12:53:21Z | |
| dc.date.issued | 2025-01-01 | |
| dc.description.abstract | We present an empirical comparison of classical baselines and pretrained transformer encoders for mental health status classification across English and Bangla social media text. Using two public datasets-an English multi-class corpus and a Bangla binary corpus-we evaluate TF-IDF with Logistic Regression and Random Forest against BERT, RoBERTa, DeBERTa, and BanglaBERT under a matched setup with a stratified eighty twenty split and macro F1 for model selection. In English, RoBERTa achieves 81.6% accuracy with a macro F1 of 78.8, while TF-IDF with Logistic Regression reaches 77.3% accuracy. In Bangla, BanglaBERT attains 88.3% accuracy with a macro F 1 of 88.3, and classical baselines surpass several non-Bangla encoders. Findings highlight the value of language models and the importance of classical machine learning models in classifying mental status across different languages. | |
| dc.description.version | Published | |
| dc.format.extent | 114-118 | |
| dc.identifier.citation | A. Islam, I. A. Rafi, S. Mondal, S. A. S. Rahman and G. R. Alam, "Which Matters More: Model or Language? An Empirical Study in English-Bangla Mental Health Classification," 2025 5th International Conference on Robotics, Automation, and Artificial Intelligence (RAAI), Singapore, Singapore, 2025, pp. 114-118, doi: 10.1109/RAAI67517.2025.11423350. | |
| dc.identifier.doi | 10.1109/RAAI67517.2025.11423350 | |
| dc.identifier.issn | 9798331558734 | |
| dc.identifier.other | 2-s2.0-105035995762 | |
| dc.identifier.uri | https://hdl.handle.net/10361/29091 | |
| dc.language.iso | en_US | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.hasversion | 10.1109/RAAI67517.2025.11423350 | |
| dc.relation.ispartof | 2025 5th International Conference on Robotics Automation and Artificial Intelligence Raai 2025 | |
| dc.relation.ispartofseries | 2025 5th International Conference on Robotics Automation and Artificial Intelligence Raai 2025 | |
| dc.relation.uri | https://ieeexplore.ieee.org/document/11423350 | |
| dc.rights | false | |
| dc.subject | Bangla | |
| dc.subject | BanglaBERT | |
| dc.subject | BERT | |
| dc.subject | Crosslingual evaluation | |
| dc.subject | DeBERTa-v3 | |
| dc.subject | Depression detection | |
| dc.subject | English | |
| dc.subject | Low-resource NLP | |
| dc.subject | Mental health text classification | |
| dc.subject | RoBERTa | |
| dc.subject | Social media | |
| dc.subject | Transformer encoders | |
| dc.subject.lcsh | Bengali language. | |
| dc.subject.lcsh | Mental health. | |
| dc.subject.lcsh | Social media. | |
| dc.title | Which matters more: Model or language? an empirical study in English-Bangla mental health classification | |
| dc.type | Conference Proceeding | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.identifier.scopus-author-id | 59710210700 | |
| person.identifier.scopus-author-id | 57567414600 | |
| person.identifier.scopus-author-id | 58161176700 | |
| person.identifier.scopus-author-id | 59710579500 | |
| person.identifier.scopus-author-id | 57348800500 |