Automated detection of unusual activities in surveillance videos remains a critical challenge due to the vast volume of footage and the rarity of anomalous events. This study proposes a novel deep learning framework that integrates 3D convolutional neural networks for spatiotemporal feature extraction, Long Short-Term Memory networks for modeling temporal dependencies, and an attention mechanism to focus on salient segments. The primary objective is to achieve highly accurate binary classification of video clips into “usual” and “unusual” categories, while addressing class imbalance and environmental variability. The model is trained and evaluated on three large-scale datasets, UCF‑Crime, XD-Violence, and CCTVFights, which comprise surveillance videos covering real-world anomalies. Experimental results show that the proposed method achieves an overall accuracy of 97.41% on the UCF-Crime dataset, 98.11% on the XD-Violence dataset, and 98.50% on the CCTVFights dataset, in addition to high precision, recall, and F1 scores on the three datasets used in the evaluation process, outperforming existing benchmarks. These findings indicate that combining spatiotemporal modeling and attention-driven context aggregation can significantly enhance anomaly detection performance in complex surveillance scenarios.
Nadia Ali (Thu,) studied this question.