Sadeque, Farig YousufSiddiqui, Md. Saiful BariKotha, Eshika EbnatChowdhury, Abtahi Bin JahangirAnan, Rafiyad KhanShowkat, Subha NajAfridi, Sayed2026-08-092026-08-0920262026-01ID 21201318ID 21201426ID 21201094ID 21201398ID 21201772https://hdl.handle.net/10361/28837This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.Cataloged from PDF version of thesis.Includes bibliographical references (pages 62-63).Automatic text summarization is a critical tool for managing the growing volume of digital content, yet effective summarization remains challenging for low-resource languages such as Bangla. This thesis investigates the capability of large language models (LLMs) to perform cross-domain Bangla text summarization under a strictly zero-shot setting. Rather than proposing a new summarization model, the study focuses on a systematic and reliable evaluation of existing models across heterogeneous domains. Summarization outputs are generated from two distinct datasets: a real-world Bangla news corpus (Prothom Alo) and the benchmark XL-Sum (Bangla) dataset. A diverse set of encoder–decoder and decoder-only LLMs is evaluated using a multi-layered assessment framework that combines traditional automatic metrics, blind LLM-as-a-Judge evaluation, SBERT-based semantic similarity analysis, and an automated error taxonomy. We assume that we need a more robust comparison beyond surface level lexical matching, which is found ineffective for Bangla abstractive summarization. However, our experimental results show that those lexical metrics (such as ROUGE and BLEU) are generally insufficient to reflect semantic quality for Bangla summary since near zero scores (‘0’scores) appear in the case of coherent summarization. In contrast, the semantic analysis shows that Banglaspecific encoder-decoder models including BanglaT5 and mT5 significantly better perform than both the multilingual and decoder-only in domains. Decoder-only models are observed to behave erratically and incline towards either ungrammatical extraction or hallucination, as is systematically verified using semantic similarity patterns and error taxonomy analysis. The findings show that fine summarization in Bangla is insensitive to surface fluency or lexical overlap but dependents on semantic abstraction and faithfulness. We believe that by presenting an exhaustive and behavior-aware evaluation framework, we are able to give practical advice for the future Bangla summarization work so as to demonstrate the importance of language-wise evaluation methodologies especially for low resource languages.74 pagesen-USAttribution-NonCommercial-NoDerivatives 4.0 InternationalBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.http://creativecommons.org/licenses/by-nc-nd/4.0/Zero-shot learningLarge language modelsLLMSText summarizationAutomatic text summarizationLow-resource languagesBangla textBengali languageError taxonomySemantic evaluationAutomatic abstracting.Natural language processing (Computer science).Text processing (Computer science).Bengali language--Data processing.Computational linguistics.Exploring cross-domain Bangla text summarization using large language modelsThesis