Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Beyond observation: the role of visual question answering in CCTV footage analysis

bracu.degree.levelUndergraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorAlam, Md. Golam Rabiul
dc.contributor.advisorReza, MD.Tanzim
dc.contributor.authorKarim, Towfiq
dc.contributor.authorRahman, MD Tashin
dc.contributor.authorRahman, MD Alvi
dc.contributor.authorUddin, Minhaj
dc.contributor.authorTabassum, Mirza Bushra
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-01-06T05:12:44Z
dc.date.available2026-01-06T05:12:44Z
dc.date.copyright2025
dc.date.issued2025-10
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 69-71).
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025.en_US
dc.description.abstractVisual Question Answering (VQA) system that revolutionizes CCTV surveillance through intelligent anomaly detection and automated incident reporting. The security and surveillance functions of CCTV cameras intensively capture terabytes of data everyday. An effective and efficient method should exist to extract footage and analyze its contents. Traditional methods find it challenging to work with highdimensional along with complex data while require long period of time and being designed for specific tasks. Deep learning models in artificial intelligence have improved research analysis functionality by making operations more efficient. Visual Question Answering (VQA) relies on Natural language processing together with computer vision to produce its operations. Our research addresses this challenge by developing an integrated framework that combines advanced computer vision with natural language processing to enable real-time, query-based video analysis and automated security reporting. An innovative CCTV surveillance system based on VQA technology and build an unified system of the combination of state-of-the-art models of computer vision, TimeSformer, UniFormer, MotionFormer, and SlowFast, and natural language processing, BLIP-2, BART, OpenCLIP, and InstructBLIP, to operate in real-time to analyze the video and provide automated feedback about the detected anomalies through the query input. TimeSformer in general and TimeSformer with spatio-temporal dynamics in particular are shown to perform better with regard to capturing spatio-temporal dynamics and are 65% accurate in determining an anomaly on our dataset. AI-powered text generation allows the system to generate rich, context-wise summaries and answers, which make it much easier to interpret and use. The experimental findings indicate that the framework is successful in the management of low-resolution video data and noisy video data, which improves the efficiency and accuracy of analysis procedure in real-time video analysis. The work offers a significant background to smart and flexible surveillance systems that can conduct proactive surveillance and accurately conduct an anomaly in the sophisticated setting.en_US
dc.description.degreeBachelor of Science in Computer Science
dc.description.statementofresponsibilityTowfiq Karim
dc.description.statementofresponsibilityMD Tashin Rahman
dc.description.statementofresponsibilityMD Alvi Rahman
dc.description.statementofresponsibilityMinhaj Uddin
dc.description.statementofresponsibilityMirza Bushra Tabassum
dc.format.extent79 pages
dc.identifier.otherID 23241106
dc.identifier.otherID 23241071
dc.identifier.otherID 24141211
dc.identifier.otherID 24141175
dc.identifier.otherID 24141174
dc.identifier.urihttp://hdl.handle.net/10361/27399
dc.language.isoenen_US
dc.publisherBRAC Universityen_US
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.subjectVQAen_US
dc.subjectVisual question answeringen_US
dc.subjectAnomaly detectionen_US
dc.subjectNatural language processingen_US
dc.subjectCCTV surveillanceen_US
dc.subjectOpenCLIPen_US
dc.subjectReal-time video analysisen_US
dc.subjectAdaptive machine learningen_US
dc.subjectMotionFormeren_US
dc.subjectUniFormeren_US
dc.subjectTimeSformeren_US
dc.subjectText generationen_US
dc.subjectBARTen_US
dc.subject.lcshComputational intelligence.
dc.subject.lcshElectronic data processing--Distributed processing.
dc.subject.lcshOptical pattern recognition.
dc.subject.lcshVideo surveillance--Real-time data processing.
dc.subject.lcshComputer vision.
dc.subject.lcshNatural language processing (Computer science).
dc.titleBeyond observation: the role of visual question answering in CCTV footage analysisen_US
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
23241106, 23241071, 24141211, 24141175, 24141174_CSE.pdf
Size:
823.73 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: