Context-aware zero-shot anomaly detection in surveillance using contrastive and predictive spatiotemporal mod
| bracu.degree.level | Undergraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Alam, Md. Ashraful | |
| dc.contributor.author | Hasan, Md. Abrar | |
| dc.contributor.author | Khan, Md. Rashid Shahriar | |
| dc.contributor.author | Justice, Mohammod Tareq Aziz | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2025-08-27T04:35:04Z | |
| dc.date.available | 2025-08-27T04:35:04Z | |
| dc.date.copyright | 2025 | |
| dc.date.issued | 2025-07 | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 41-43). | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025. | en_US |
| dc.description.abstract | Tackling anomalies through surveillance feeds is challenging due to the unpredictable nature of anomalies and their strong dependence on context. Modern video anomaly detection architectures have been shown to thrive in such conditions. Their ability to adapt to intricate and complex patterns serves as the foundation of anomaly detection, especially for unseen scenarios, making the impossible seem tangible. The research demonstrates a novel context-aware zero-shot anomaly detection framework that learns normal spatiotemporal patterns and identifies anomalies without any explicit anomaly examples during training. In order to perform the approach, it proposed a hybrid model which is a combination of TimeSformer, DPC, and CLIP. A TimeSformer-based vision transformer backbone is employed to encode video sequences, capturing rich spatial-temporal features. We integrate Data Predictive Control (DPC) to forecast future video dynamics and flag deviations. Simultaneously, we leverage the vision-language power of CLIP in a semantic stream where the model is conditioned on contextual information and uses text prompts to detect concept-level irregularities in a zero-shot fashion. These components are jointly optimized using InfoNCE and Contrastive Predictive Coding (CPC) losses, enabling the model to align video inputs with their semantic and temporal contexts without ever being exposed to anomaly labels. To condition decisions on the situational context, we propose a context-gating mechanism that modulates temporal predictions based on scene-specific text or global video features. During inference, anomalies are flagged based on a fusion of context misalignment and predictive failure, allowing the system to generalize to previously unseen behaviours. Evaluations of our lightweight and fully zero-shot approach achieve a ROC-AUC of 84.5 %, and a PR-AUC of 72.3 %. This work advances the gap between semantic understanding and temporal prediction in surveillance, laying the foundation for context-sensitive, zero-shot detection systems deployable in dynamic real-world environments. | en_US |
| dc.description.degree | Bachelor of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Md. Abrar Hasan | |
| dc.description.statementofresponsibility | Md. Rashid Shahriar Khan | |
| dc.description.statementofresponsibility | Mohammod Tareq Aziz Justice | |
| dc.format.extent | 43 pages | |
| dc.identifier.other | ID 23241115 | |
| dc.identifier.other | ID 20101557 | |
| dc.identifier.other | ID 21201585 | |
| dc.identifier.uri | http://hdl.handle.net/10361/26592 | |
| dc.language.iso | en | en_US |
| dc.publisher | BRAC University | en_US |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.subject | Context awareness | en_US |
| dc.subject | Zero-shot learning | en_US |
| dc.subject | Video surveillance | en_US |
| dc.subject | Contrastive learning | en_US |
| dc.subject | Spatiotemporal | en_US |
| dc.subject | Predictive modeling | en_US |
| dc.subject.lcsh | Data Visualization. | |
| dc.subject.lcsh | Electronic surveillance. | |
| dc.subject.lcsh | Spatial analysis. | |
| dc.subject.lcsh | Geographic information systems. | |
| dc.subject.lcsh | Prediction theory. | |
| dc.subject.lcsh | Mathematical models. | |
| dc.title | Context-aware zero-shot anomaly detection in surveillance using contrastive and predictive spatiotemporal mod | en_US |
| dc.type | Thesis | en_US |