Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Video Anomaly Detection (VAD) has become an essential component of modern surveillance systems, particularly because the number of cameras in public spaces produces a massive amount of video data that cannot be manually monitored in real time. Traditional approaches to VAD largely focus on spatiotemporal features to identify irregular patterns; however, these models have two major limitations: lack of contextual awareness of the surrounding scene and difficulty in modeling temporal dependencies between video snippets. As a result, these models frequently misclassify normal context activities as anomalies or fail to detect subtle abnormal events. To address these issues, an integrated framework, contrast-learning for scene embedding, and a transformer encoder with Multi Instance Learning (MIL)-guided Attention Mechanism (STAM) in an MIL setup are proposed. Here, Scene embeddings are obtained using contrastive learning to capture environment-specific contexts, while Inflated 3D ConvNet (13D) captured spatiotemporal features are modeled through a transformer encoder that uses MIL attention to prioritize anomalous snippets. From the results, the evaluation on the benchmark dataset demonstrated that the proposed STAM model achieved superior accuracy (99.85%) and an Area Under Cure (AUC) of (99.98%) when compared to the existing Dual Network in MIL (DD-MIL).</p>

Show More

Keywords

video scene spatiotemporal features models

Related Articles

PORE

About

Connect