Beyond observation: the role of visual question answering in CCTV footage analysis
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Alam, Md. Golam Rabiul | |
| dc.contributor.advisor | Reza, MD.Tanzim | |
| dc.contributor.author | Karim, Towfiq | |
| dc.contributor.author | Rahman, MD Tashin | |
| dc.contributor.author | Rahman, MD Alvi | |
| dc.contributor.author | Uddin, Minhaj | |
| dc.contributor.author | Tabassum, Mirza Bushra | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-01-06T05:12:44Z | |
| dc.date.available | 2026-01-06T05:12:44Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-10 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 69-71). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025. | en_US |
| dc.description.abstract | Visual Question Answering (VQA) system that revolutionizes CCTV surveillance through intelligent anomaly detection and automated incident reporting. The security and surveillance functions of CCTV cameras intensively capture terabytes of data everyday. An effective and efficient method should exist to extract footage and analyze its contents. Traditional methods find it challenging to work with highdimensional along with complex data while require long period of time and being designed for specific tasks. Deep learning models in artificial intelligence have improved research analysis functionality by making operations more efficient. Visual Question Answering (VQA) relies on Natural language processing together with computer vision to produce its operations. Our research addresses this challenge by developing an integrated framework that combines advanced computer vision with natural language processing to enable real-time, query-based video analysis and automated security reporting. An innovative CCTV surveillance system based on VQA technology and build an unified system of the combination of state-of-the-art models of computer vision, TimeSformer, UniFormer, MotionFormer, and SlowFast, and natural language processing, BLIP-2, BART, OpenCLIP, and InstructBLIP, to operate in real-time to analyze the video and provide automated feedback about the detected anomalies through the query input. TimeSformer in general and TimeSformer with spatio-temporal dynamics in particular are shown to perform better with regard to capturing spatio-temporal dynamics and are 65% accurate in determining an anomaly on our dataset. AI-powered text generation allows the system to generate rich, context-wise summaries and answers, which make it much easier to interpret and use. The experimental findings indicate that the framework is successful in the management of low-resolution video data and noisy video data, which improves the efficiency and accuracy of analysis procedure in real-time video analysis. The work offers a significant background to smart and flexible surveillance systems that can conduct proactive surveillance and accurately conduct an anomaly in the sophisticated setting. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science | |
| dc.description.statementofresponsibility | Towfiq Karim | |
| dc.description.statementofresponsibility | MD Tashin Rahman | |
| dc.description.statementofresponsibility | MD Alvi Rahman | |
| dc.description.statementofresponsibility | Minhaj Uddin | |
| dc.description.statementofresponsibility | Mirza Bushra Tabassum | |
| dc.format.extent | 79 pages | |
| dc.identifier.other | ID 23241106 | |
| dc.identifier.other | ID 23241071 | |
| dc.identifier.other | ID 24141211 | |
| dc.identifier.other | ID 24141175 | |
| dc.identifier.other | ID 24141174 | |
| dc.identifier.uri | http://hdl.handle.net/10361/27399 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | VQA | en_US |
| dc.subject | Visual question answering | en_US |
| dc.subject | Anomaly detection | en_US |
| dc.subject | Natural language processing | en_US |
| dc.subject | CCTV surveillance | en_US |
| dc.subject | OpenCLIP | en_US |
| dc.subject | Real-time video analysis | en_US |
| dc.subject | Adaptive machine learning | en_US |
| dc.subject | MotionFormer | en_US |
| dc.subject | UniFormer | en_US |
| dc.subject | TimeSformer | en_US |
| dc.subject | Text generation | en_US |
| dc.subject | BART | en_US |
| dc.subject.lcsh | Computational intelligence. | |
| dc.subject.lcsh | Electronic data processing--Distributed processing. | |
| dc.subject.lcsh | Optical pattern recognition. | |
| dc.subject.lcsh | Video surveillance--Real-time data processing. | |
| dc.subject.lcsh | Computer vision. | |
| dc.subject.lcsh | Natural language processing (Computer science). | |
| dc.title | Beyond observation: the role of visual question answering in CCTV footage analysis | en_US |
| dc.type | Thesis | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 23241106, 23241071, 24141211, 24141175, 24141174_CSE.pdf
- Size:
- 823.73 KB
- Format:
- Adobe Portable Document Format
- Description:
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: