Introducing a Bangla sentence gloss pair dataset for Bangla sign language translation and research
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Rasel, Annajiat Alim | |
| dc.contributor.author | Roudra, Nafis Ashraf | |
| dc.contributor.author | Saha, Neelavro | |
| dc.contributor.author | Shahriyar, Rafi | |
| dc.contributor.author | Sakib, Saadman | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2025-08-31T06:39:41Z | |
| dc.date.available | 2025-08-31T06:39:41Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-06 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 42-44). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025. | en_US |
| dc.description.abstract | Bangla Sign Language translation and recognition has been an evolving research topic throughout the years. However, existing research on this field is limited to word and alphabet level detection. For a more continuous sentence level detection of spoken Bangla sentences and their corresponding signed gestures, structured and comprehensive annotations are essential. Therefore, in this paper we introduce a dataset that consists of Bangla sentences and their matching gloss sequence pairs. Gloss sequences are made up of individual glosses which are Bangla sign supported words and serve as an intermediate representation for a continuous sign. Our dataset consists of 1000 high quality Bangla sentences that are manually annotated into a gloss sequence by a professional signer. With no available open-source datasets on Bangla sentence and gloss pairs, we additionally augment our dataset with 3000 synthetic samples. For the augmentation process, we introduce a RAG-based pipeline which incorporates rule-based linguistic strategies and prompt engineering techniques that we have adopted by critically analyzing our human annotated sentencegloss pairs and by working closely with our professional signer. Furthermore, we finetune several transformer-based models such as mBart-50, Google mT5, GPT4.1- nano and perform BLEU score based evaluations to determine which model performs the best in the task of Sentence-to-gloss translation. Finally, based on these evaluation metrics we also compare how our dataset performs against the Phoenix-2014T dataset. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Nafis Ashraf Roudra | |
| dc.description.statementofresponsibility | Neelavro Saha | |
| dc.description.statementofresponsibility | Rafi Shahriyar | |
| dc.description.statementofresponsibility | Saadman Sakib | |
| dc.format.extent | 45 pages | |
| dc.identifier.other | ID 21301410 | |
| dc.identifier.other | ID 21301181 | |
| dc.identifier.other | ID 21301198 | |
| dc.identifier.other | ID 21101091 | |
| dc.identifier.uri | http://hdl.handle.net/10361/26615 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Natural language processing | en_US |
| dc.subject | Transformers | en_US |
| dc.subject | Machine translation | en_US |
| dc.subject | Bangla sign language | en_US |
| dc.subject | Data augmentation | en_US |
| dc.subject.lcsh | Electric transformers. | |
| dc.subject.lcsh | Natural language processing (Computer science). | |
| dc.subject.lcsh | Machine learning. | |
| dc.subject.lcsh | Data mining. | |
| dc.title | Introducing a Bangla sentence gloss pair dataset for Bangla sign language translation and research | en_US |
| dc.type | Thesis | en_US |