Detecting derogatory comments on women using transformer-based models

Citation

S. J. Prithila et al., "Detecting Derogatory Comments on Women using Transformer-Based Models," 2023 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT), Malang, Indonesia, 2023, pp. 278-284, doi: 10.1109/COMNETSAT59769.2023.10420692.

Abstract

Natural Language Processing (NLP) is a piqued interest field nowadays, as it helps AI to understand and interpret human languages. In order to facilitate the advancement in this field, in this paper, we propose research on the detection of derogatory comments against women with the help of transformer-based models. Here, our main focus is to detect misogynistic comments, as the women of our country mainly get harassed by such texts. This paper aims to make a comparative study on how efficient transformer models are in detecting gender-biased slandering in languages such as English and Bengali. To carry out this research procedure, the datasets we used were in English and Bengali languages which were further trained across the following transformer models: BanglaBERT, XLM-RoBERTa, m-BERT, and DistilBERT. To give further richness to the paper, the Bengali and English datasets used were created by combining multiple different datasets in these languages. The datasets were extracted from various papers related to this or a similar field of research to help reduce biases and improve language understanding capability. Upon, training our datasets across the mentioned models, for the Bengali dataset, Bangla-BERT-Base performed the best with an F1 score of 94% and for the English dataset, m-BERT scored the best with an F1 score of 86.1%. To add on, since the paper mostly focuses on the Bengali language, it will furthermore, encourage others to increase research on low-resourced languages.

Description

Type

Conference Proceeding