Bangla SBERT - Sentence embedding using multilingual knowledge distillation

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorUddin M.S.
dc.contributor.authorHaque M.A.
dc.contributor.authorRifat R.H.
dc.contributor.authorKamal, Marufa
dc.contributor.authorGupta K.D.
dc.contributor.authorGeorge R.
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-13T08:15:16Z
dc.date.available2026-09-13T08:15:16Z
dc.date.issued2024-01-01
dc.description.abstractWord embeddings have revolutionized NLP by capturing the semantic associations between words effectively. However, sentence embeddings present notable benefits in the context of advanced language comprehension. While these developments have impacted high-resource languages like English, low-resource languages like Bangla have not benefited as much. The objective of this study is to address this disparity by creating sentence embeddings for the Bangla language, to improve information retrieval, sentiment analysis, content suggestion, etc. In this study, we developed Bangla Sentence-BERT by fine-tuning it on novel datasets generated through machine translation and utilizing diverse open-source datasets. The approach we used consisted of utilizing the stsb-xlm-r-multilingual as the teacher model and XLM-RoBERTa (XLMR) as the student model for multilingual interpretation. We evaluated the efficacy of our suggested approach on multilingual sentence-BERT models and classical machine learning algorithms. The performance of our model was remarkable as it achieved an accuracy of 97 % on real text classification. The results demonstrate the efficacy of our Bangla sentence transformer model in comprehending meaning and its potential for a range of Bangla natural language processing applications, such as text classification.
dc.description.versionPublished
dc.format.extent495-500
dc.identifier.citationM. S. Uddin, M. A. Haque, R. H. Rifat, M. Kamal, K. D. Gupta and R. George, "Bangla SBERT - Sentence Embedding Using Multilingual Knowledge Distillation," 2024 IEEE 15th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON), Yorktown Heights, NY, USA, 2024, pp. 495-500, doi: 10.1109/UEMCON62879.2024.10754765.
dc.identifier.doi10.1109/UEMCON62879.2024.10754765
dc.identifier.issn9798331540906
dc.identifier.other2-s2.0-85212669599
dc.identifier.urihttps://hdl.handle.net/10361/29870
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/UEMCON62879.2024.10754765
dc.relation.ispartof2024 IEEE 15th Annual Ubiquitous Computing Electronics and Mobile Communication Conference Uemcon 2024
dc.relation.ispartofseries2024 IEEE 15th Annual Ubiquitous Computing Electronics and Mobile Communication Conference Uemcon 2024
dc.relation.urihttps://ieeexplore.ieee.org/document/10754765
dc.subjectBangla NLP
dc.subjectKnowledge distillation
dc.subjectSBERT
dc.subjectSentence similarity
dc.subjectSentence transformer
dc.subject.lcshComputational linguistics.
dc.subject.lcshDistillation.
dc.titleBangla SBERT - Sentence embedding using multilingual knowledge distillation
dc.typeConference Proceeding
person.affiliation.nameComilla University
person.affiliation.nameClark Atlanta University
person.affiliation.nameEdward E. Whitacre Jr. College of Engineering
person.affiliation.nameBRAC University
person.affiliation.nameClark Atlanta University
person.affiliation.nameClark Atlanta University
person.identifier.scopus-author-id59212858900
person.identifier.scopus-author-id57219243705
person.identifier.scopus-author-id58306614600
person.identifier.scopus-author-id58170084700
person.identifier.scopus-author-id57205211355
person.identifier.scopus-author-id7402637244

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Demo.jpg
Size:
27.28 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: