Mukta, Jannatun NoorHossain, AriyanNabi, Syed MahbubunAnik, Al Jami IslamMazlish, Ahsan Shariar KhanAhsan, Mustafis2026-08-062026-08-0620262026-01ID 22101472ID 22101577ID 22101635ID 22101486https://hdl.handle.net/10361/28820This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2026.Cataloged from PDF version of thesis.Includes bibliographical references (pages 51-53).Interpreting the semantic roles of entities in memes requires navigating complex interactions between visual images, embedded texts, and implicit cultural knowledge. This is more challenging for languages with limited resources such as Bengali. Most existing works on this topic have only analyzed general concepts such as hate speech detection and sentiment analysis. They have not considered the detailed narrative roles that are required to comprehend the meaning of the memes. Our goal is to fill this gap and create an entity-based framework for semantic role labeling and humor explanation. To accomplish this, we have developed a labeled dataset of 2,658 Bengali memes, each labeled with roles such as Hero, Villain, and Victim, along with local context notes that were used to resolve cultural misunderstandings during training. We propose and evaluate two complementary frameworks: a Triple-Stream Fusion model utilizing cross-attention, and a Vision-Language Model (VLM) pipeline leveraging Qwen2-VL for holistic reasoning. Experimental results demonstrate that the VLM approach significantly outperforms the fusion baseline, achieving an accuracy of 0.8068 and a Macro-F1 score of 0.7953. Furthermore, our fine-tuned humor explanation generator achieved a significantly improved LLM-asa- Judge quality score of 7.33/10 (compared to a 6.03 baseline), confirming its ability to produce culturally grounded interpretations. Our study is a strong benchmark on the task of examining satire and social commentary in Bengali memes and provides a simple approach to jointly label roles and generate explanations by combining images, text, and context.63 pagesen-USAttribution-NonCommercial-NoDerivatives 4.0 InternationalBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.http://creativecommons.org/licenses/by-nc-nd/4.0/Bangla memesContextualizationComputer visionContext-aware modelMultimodal communicationMachine learningInternet cultureInternet memesSemantic analysisComputational linguisticsNatural language processing (Computer science).Memes--Bangladesh.Bengali wit and humor.Linguistic analysis (Linguistics).Bengali language--Data processing.Analyzing the semantic roles in Bangladeshi memes: A multimodal approach to contextual role labeling and explanation generationThesis