Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Beyond observation: the role of visual question answering in CCTV footage analysis

Citation

Abstract

Visual Question Answering (VQA) system that revolutionizes CCTV surveillance through intelligent anomaly detection and automated incident reporting. The security and surveillance functions of CCTV cameras intensively capture terabytes of data everyday. An effective and efficient method should exist to extract footage and analyze its contents. Traditional methods find it challenging to work with highdimensional along with complex data while require long period of time and being designed for specific tasks. Deep learning models in artificial intelligence have improved research analysis functionality by making operations more efficient. Visual Question Answering (VQA) relies on Natural language processing together with computer vision to produce its operations. Our research addresses this challenge by developing an integrated framework that combines advanced computer vision with natural language processing to enable real-time, query-based video analysis and automated security reporting. An innovative CCTV surveillance system based on VQA technology and build an unified system of the combination of state-of-the-art models of computer vision, TimeSformer, UniFormer, MotionFormer, and SlowFast, and natural language processing, BLIP-2, BART, OpenCLIP, and InstructBLIP, to operate in real-time to analyze the video and provide automated feedback about the detected anomalies through the query input. TimeSformer in general and TimeSformer with spatio-temporal dynamics in particular are shown to perform better with regard to capturing spatio-temporal dynamics and are 65% accurate in determining an anomaly on our dataset. AI-powered text generation allows the system to generate rich, context-wise summaries and answers, which make it much easier to interpret and use. The experimental findings indicate that the framework is successful in the management of low-resolution video data and noisy video data, which improves the efficiency and accuracy of analysis procedure in real-time video analysis. The work offers a significant background to smart and flexible surveillance systems that can conduct proactive surveillance and accurately conduct an anomaly in the sophisticated setting.

Description

Cataloged from PDF version of thesis.
Includes bibliographical references (pages 69-71).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025.

Publisher Link

Type

Thesis