Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Real-time scene description and interpretation using zero-shot learning and prompt-engineered vision-language models

bracu.degree.levelPostgraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorAnwar, Md. Tawhid
dc.contributor.authorApurba, Md Shifatul Ahsan
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-04-08T04:42:19Z
dc.date.available2026-04-08T04:42:19Z
dc.date.copyright2025
dc.date.issued2024-08
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 32-34).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science, 2025.en_US
dc.description.abstractReal-time scene description and interpretation are essential for diverse applications such as surveillance, interactive media, and automated video analysis. However, most existing methods rely heavily on large-scale labeled datasets, thereby limiting their adaptability in dynamic or previously unseen scenarios. In this work, we propose a novel mixed-model framework that integrates Vision-Language Models (VLMs), Large Language Models (LLMs), and lightweight object detection networks (e.g., MobileNet-SSD) through advanced prompt engineering. By leveraging zeroshot learning, our approach generates contextually rich scene descriptions without requiring domain-specific or task-specific retraining. The prompt engineering component reduces sensitivity to subtle linguistic variations, enhancing robustness across diverse input formulations. Furthermore, the lightweight detector ensures real-time performance, making the framework suitable for resource-constrained environments. To address ethical and fairness considerations, we incorporate bias mitigation strategies that limit the propagation of harmful stereotypes from large-scale pretraining data. Experimental evaluations on multiple open-domain scenarios demonstrate that our system offers reliable and efficient scene interpretation, maintaining high accuracy in challenging conditions where traditional supervised techniques often fail. This research paves the way for more flexible, scalable, and responsible visionlanguage systems capable of operating effectively in real-world, zero-shot contexts.en_US
dc.description.degreeMaster of Science in Computer Science
dc.description.statementofresponsibilityMd Shifatul Ahsan Apurba
dc.format.extent34 pages
dc.identifier.otherID 22266027
dc.identifier.urihttp://hdl.handle.net/10361/27808
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectVision-Language Modelsen_US
dc.subjectVLMen_US
dc.subjectLarge Language Modelsen_US
dc.subjectLLMen_US
dc.subjectZero-Shot Learningen_US
dc.subjectMachine learningen_US
dc.subject.lcshReal-time data processing.
dc.subject.lcshComputer vision.
dc.subject.lcshNatural language processing (Computer science).
dc.subject.lcshImage processing--Digital techniques--Data processing.
dc.subject.lcshDigital video.
dc.titleReal-time scene description and interpretation using zero-shot learning and prompt-engineered vision-language modelsen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
22266027_CSE.pdf
Size:
781.63 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: