Context-aware zero-shot anomaly detection in surveillance using contrastive and predictive spatiotemporal mod
Loading...
Date
Publisher
BRAC University
Citation
Abstract
Tackling anomalies through surveillance feeds is challenging due to the unpredictable
nature of anomalies and their strong dependence on context. Modern video anomaly
detection architectures have been shown to thrive in such conditions. Their ability
to adapt to intricate and complex patterns serves as the foundation of anomaly detection,
especially for unseen scenarios, making the impossible seem tangible. The
research demonstrates a novel context-aware zero-shot anomaly detection framework
that learns normal spatiotemporal patterns and identifies anomalies without
any explicit anomaly examples during training. In order to perform the approach, it
proposed a hybrid model which is a combination of TimeSformer, DPC, and CLIP.
A TimeSformer-based vision transformer backbone is employed to encode video sequences,
capturing rich spatial-temporal features. We integrate Data Predictive
Control (DPC) to forecast future video dynamics and flag deviations. Simultaneously,
we leverage the vision-language power of CLIP in a semantic stream where
the model is conditioned on contextual information and uses text prompts to detect
concept-level irregularities in a zero-shot fashion. These components are jointly optimized
using InfoNCE and Contrastive Predictive Coding (CPC) losses, enabling
the model to align video inputs with their semantic and temporal contexts without
ever being exposed to anomaly labels. To condition decisions on the situational context,
we propose a context-gating mechanism that modulates temporal predictions
based on scene-specific text or global video features. During inference, anomalies
are flagged based on a fusion of context misalignment and predictive failure, allowing
the system to generalize to previously unseen behaviours. Evaluations of
our lightweight and fully zero-shot approach achieve a ROC-AUC of 84.5 %, and a
PR-AUC of 72.3 %. This work advances the gap between semantic understanding
and temporal prediction in surveillance, laying the foundation for context-sensitive,
zero-shot detection systems deployable in dynamic real-world environments.
Description
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 41-43).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
Includes bibliographical references (pages 41-43).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science and Engineering, 2025.
Publisher Link
Type
Thesis