End-to-end pipeline: Normalization, and summarization of Bangla-English code-switching conversation
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Islam, Nazmul | |
| dc.contributor.author | Rahman, Samir | |
| dc.contributor.author | Siddique, Dania | |
| dc.contributor.author | Tasnim, Humaira Sadia | |
| dc.contributor.author | Khan, Zahidul Islam | |
| dc.contributor.author | Omar, Nayem Bin | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-08-16T08:13:43Z | |
| dc.date.available | 2026-08-16T08:13:43Z | |
| dc.date.copyright | 2026 | |
| dc.date.issued | 2026-01 | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026. | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 57-59). | |
| dc.description.abstract | Conventional Natural Language Processing (NLP) systems are predominantly designed and trained for monolingual text. However, the extensive use of Bangla-English and Banglish( Bengali written in Romanized alphabets) code-switching informal conversations in digital communication proves to be challenging for these NLP systems. To addresses this research gap in processing mixed language texts of Bangla-English-Banglish we proposed BiLoRA-BN, a end-to-end pipeline specifically designed for normalization and summarization of code-switched Bangla-English-Banglish conversations in digital communication. The proposed system employs a two-stage Low-Rank Adaptation (LoRA) architecture built on a shared, pre-trained Transformer as backbone with 4-bit quantization, enabling efficient multi-task learning while reducing trainable parameters. The experimental results of BiLoRA-BN are compelling, it significantly outperforms conventional sequential and cascading pipelines, with a +5.74 BLEU gain in normalization quality and a +5.26 ROUGE-1 improvement in final summary accuracy. During interface testing BiLoRA-BN also delivers results faster compared to other pipeline based cascading approach of different architecture and pre-trained models. Crucially, in the interface part, the entire system of BiLoRA-BN can operates on a consumer-grade GPUs with 8GB of memory. By directly modeling all the transformations BiLoRA-BN tries to capture the reality of multilingual digital discourse, with the complex scenario like code-switching and code-mixing in the conversations. This work contributes a step toward understanding how people naturally speak and write to communicate in digital spaces and how NLP model work with it. | |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Samir Rahman | |
| dc.description.statementofresponsibility | Dania Siddique | |
| dc.description.statementofresponsibility | Humaira Sadia Tasnim | |
| dc.description.statementofresponsibility | Zahidul Islam Khan | |
| dc.description.statementofresponsibility | Nayem Bin Omar | |
| dc.format.extent | 69 pages | |
| dc.identifier.other | ID 24341192 | |
| dc.identifier.other | ID 24241088 | |
| dc.identifier.other | ID 24241055 | |
| dc.identifier.other | ID 20301158 | |
| dc.identifier.other | ID 20301435 | |
| dc.identifier.uri | https://hdl.handle.net/10361/29147 | |
| dc.language.iso | en_US | |
| dc.publisher | BRAC University | |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | en |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject | Natural language processing | |
| dc.subject | Code-switching | |
| dc.subject | Text normalization | |
| dc.subject | Abstractive summarization | |
| dc.subject | Low-rank adaptation | |
| dc.subject | LoRA | |
| dc.subject | Multilingual models | |
| dc.subject | Quantization | |
| dc.subject | Computational linguistics | |
| dc.subject | Artificial intelligence | |
| dc.subject.lcsh | Natural language processing (Computer science). | |
| dc.subject.lcsh | Code switching (Linguistics). | |
| dc.subject.lcsh | Text processing (Computer science). | |
| dc.subject.lcsh | Automatic abstracting. | |
| dc.subject.lcsh | Machine translating. | |
| dc.subject.lcsh | Translating and interpreting. | |
| dc.subject.lcsh | Bengali language--Translating. | |
| dc.title | End-to-end pipeline: Normalization, and summarization of Bangla-English code-switching conversation | |
| dc.type | Thesis |