Alam, Md. Golam RabiulReza, MD.TanzimKarim, TowfiqRahman, MD TashinRahman, MD AlviUddin, MinhajTabassum, Mirza Bushra2026-01-062026-01-0620252025-10ID 23241106ID 23241071ID 24141211ID 24141175ID 24141174http://hdl.handle.net/10361/27399Cataloged from PDF version of thesis.Includes bibliographical references (pages 69-71).This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2025.Visual Question Answering (VQA) system that revolutionizes CCTV surveillance through intelligent anomaly detection and automated incident reporting. The security and surveillance functions of CCTV cameras intensively capture terabytes of data everyday. An effective and efficient method should exist to extract footage and analyze its contents. Traditional methods find it challenging to work with highdimensional along with complex data while require long period of time and being designed for specific tasks. Deep learning models in artificial intelligence have improved research analysis functionality by making operations more efficient. Visual Question Answering (VQA) relies on Natural language processing together with computer vision to produce its operations. Our research addresses this challenge by developing an integrated framework that combines advanced computer vision with natural language processing to enable real-time, query-based video analysis and automated security reporting. An innovative CCTV surveillance system based on VQA technology and build an unified system of the combination of state-of-the-art models of computer vision, TimeSformer, UniFormer, MotionFormer, and SlowFast, and natural language processing, BLIP-2, BART, OpenCLIP, and InstructBLIP, to operate in real-time to analyze the video and provide automated feedback about the detected anomalies through the query input. TimeSformer in general and TimeSformer with spatio-temporal dynamics in particular are shown to perform better with regard to capturing spatio-temporal dynamics and are 65% accurate in determining an anomaly on our dataset. AI-powered text generation allows the system to generate rich, context-wise summaries and answers, which make it much easier to interpret and use. The experimental findings indicate that the framework is successful in the management of low-resolution video data and noisy video data, which improves the efficiency and accuracy of analysis procedure in real-time video analysis. The work offers a significant background to smart and flexible surveillance systems that can conduct proactive surveillance and accurately conduct an anomaly in the sophisticated setting.79 pagesenBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.VQAVisual question answeringAnomaly detectionNatural language processingCCTV surveillanceOpenCLIPReal-time video analysisAdaptive machine learningMotionFormerUniFormerTimeSformerText generationBARTComputational intelligence.Electronic data processing--Distributed processing.Optical pattern recognition.Video surveillance--Real-time data processing.Computer vision.Natural language processing (Computer science).Beyond observation: the role of visual question answering in CCTV footage analysisThesis