Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Context-aware zero-shot anomaly detection in surveillance using contrastive and predictive spatiotemporal mod

Citation

Abstract

Tackling anomalies through surveillance feeds is challenging due to the unpredictable nature of anomalies and their strong dependence on context. Modern video anomaly detection architectures have been shown to thrive in such conditions. Their ability to adapt to intricate and complex patterns serves as the foundation of anomaly detection, especially for unseen scenarios, making the impossible seem tangible. The research demonstrates a novel context-aware zero-shot anomaly detection framework that learns normal spatiotemporal patterns and identifies anomalies without any explicit anomaly examples during training. In order to perform the approach, it proposed a hybrid model which is a combination of TimeSformer, DPC, and CLIP. A TimeSformer-based vision transformer backbone is employed to encode video sequences, capturing rich spatial-temporal features. We integrate Data Predictive Control (DPC) to forecast future video dynamics and flag deviations. Simultaneously, we leverage the vision-language power of CLIP in a semantic stream where the model is conditioned on contextual information and uses text prompts to detect concept-level irregularities in a zero-shot fashion. These components are jointly optimized using InfoNCE and Contrastive Predictive Coding (CPC) losses, enabling the model to align video inputs with their semantic and temporal contexts without ever being exposed to anomaly labels. To condition decisions on the situational context, we propose a context-gating mechanism that modulates temporal predictions based on scene-specific text or global video features. During inference, anomalies are flagged based on a fusion of context misalignment and predictive failure, allowing the system to generalize to previously unseen behaviours. Evaluations of our lightweight and fully zero-shot approach achieve a ROC-AUC of 84.5 %, and a PR-AUC of 72.3 %. This work advances the gap between semantic understanding and temporal prediction in surveillance, laying the foundation for context-sensitive, zero-shot detection systems deployable in dynamic real-world environments.

Description

Cataloged from PDF version of thesis.
Includes bibliographical references (pages 41-43).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.

Publisher Link

Type

Thesis