Video anomaly detection aims to identify rare, unexpected events in real-world surveillance environments, where the diversity and scarcity of annotated examples make exhaustive supervision impractical. Prediction-based unsupervised learning addresses this problem by modeling normal spatiotemporal patterns and flagging anomalies through prediction deviations. However, existing approaches exhibit limited spatial representation capability and insufficient understanding of clip-level temporal dependency due to appearance variations and complex motion dynamics within the scene. To address these challenges, this paper introduces RMTA-Net, a future-frame prediction network for anomaly detection that jointly learns spatial representations and temporally coherent dependencies of sequential video frames. First, a residual spatial feature enhancement network progressively extracts structural and appearance information at multiple levels. The independently encoded spatial features are organized as an explicit temporal representation. Subsequently, a recurrent conditioned memory-guided temporal attention (RMTA) module integrates recurrent temporal processing with learnable memory banks and memory-guided temporal attention to model inter-frame dependencies within the observed clip. The dual-branch pipeline simultaneously processes the recurrent network output using attention-based memory access and temporal attention-driven global normality-prior retrieval. The learnable memory banks encode dataset-level normality priors, and adaptive gating fuses retrieved recurrent memory with contextual temporal relationships. Finally, an attention-enhanced decoder predicts future frames from the learned spatiotemporal embeddings, where anomalies are identified using prediction discrepancies. Extensive experiments were conducted on three benchmark datasets, namely UCSD Ped2, CUHK Avenue, and ShanghaiTech, demonstrating the effectiveness of the proposed method. RMTA-Net achieved frame-level AUC scores of 99.0%, 90.1%, and 75.8%, respectively, and remained competitive with several state-of-the-art methods.
No takes yet. Share an insight, caveat, or question.
Chouhan et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: