PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2026ACM Transactions on Multimedia Computing Communications and Applications3 citations

A New Semi-Supervised Video Anomaly Detection Baseline in Lack of Anomalous Samples

View Full Paper
MZMengyang ZhaoHYHaiyang YuTFTeng Fu

Key Points

  • The aim is to improve video anomaly detection in situations with limited anomaly samples by re-framing the task.
  • Redefine video anomaly detection as out-of-distribution detection.
  • Leverage large visual language models to generate text summaries and extract visual features.
  • Combine language features from text encoder with visual features for robust multimodal representation.
  • Implement detection method to learn normality centers from normal and unlabeled samples.
  • The proposed method shows improved robustness with fewer anomalous classes and samples.
  • Experiments indicate effective anomaly detection, even with limited available data.

Abstract

Video anomaly detection (VAD) has been widely studied for its important applications in multimedia community. Recently, many Weakly-Supervised VAD (WS-VAD) methods have been proposed, which tend to treat VAD as a classification task through multiple instance learning and result in the need to collect sufficient anomaly classes and samples to be used for training a classifier. However, anomaly events tend to be open-set and rare in real-world applications, so we often have difficulty collecting all anomaly classes and enough sample anomalies, which is a difficult situation for WS-VAD to cope with. To this end, we consider to treat VAD as an out-of-distribution detection task rather than a classification task and propose a simple but effective semi-supervised baseline method. First, we leverage the powerful zero-shot capability of large visual language models to generate summary text descriptions for videos and extract visual features as intermediates for subsequent use. Next, we use a text encoder to extract language features and combine them with visual features to obtain robust multimodal features. Finally, we introduce an out-of-distribution detection method learns the center of normality in multimodal space from normal and unlabeled samples, while deviating abnormal samples from the center to cope with the scarcity of abnormal samples. To implement our baseline method, we also provide a new semi-supervised dataset by reorganizing an existing benchmark, which is the first available dataset in the VAD community that provides trimmed videos consisting of complete abnormal events. Experiments demonstrate that our method performs more robustly when fewer anomaly classes and anomaly samples collected.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhao et al. (2026) studied this question.

synapsesocial.com/papers/69a287570a974eb0d3c03023https://doi.org/10.1145/3797034
Ask AI
Helpful
Bookmark
Share
View Full Paper