PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 10, 2026Sensors0 citationsOpen Access

Learning Fine-Grained Video Anomaly Detection from Normal Videos

View Full Paper
RWRuqin WangYTYasumasa TamuraMYMasahito Yamamoto

Key Points

  • This research aims to improve video anomaly detection by generating fine-grained annotations from normal videos.
  • Developed a framework for unsupervised video anomaly generation using only normal videos.
  • Leveraged VLMs to create structured textual descriptions of potential anomalies.
  • Constructed a fine-grained VAD network to produce video-level, frame-level, and region-level predictions.
  • Achieved significant improvements in fine-grained anomaly detection accuracy compared to previous methods.
  • Successfully generated abnormal segments with high realism from normal video inputs.
  • Demonstrated the capability to provide detailed annotations at multiple levels.

Abstract

Video anomaly detection (VAD) aims to identify abnormal events in videos. Due to the lack of high-quality training data with detailed annotations, current VAD methods can only produce video-level predictions. To remedy this, several methods attempt to synthesize pseudo video anomalies. However, these methods suffer from low realism and coarse annotations, which limits their performance in real-world scenarios. In this paper, we propose a framework for unsupervised anomaly video generation from solely normal videos, leveraging VLMs to generate structured textual descriptions of anomalies conditioned on the perception of this video. Then, abnormal segments are synthesized using VLMs based on the synthetic textual descriptions. As our framework is highly controllable, video-level and region-level labels can be obtained to provide fine-grained annotations. On top of the synthetic data, we develop a fine-grained VAD network to simultaneously produce video-level, frame-level, and region-level predictions. Experiments show that our method achieves remarkable fine-grained VAD performance.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2026) studied this question.

synapsesocial.com/papers/6a508bde6eeac72a437a0409https://doi.org/10.3390/s26134314
Ask AI
Helpful
Bookmark
Share
View Full Paper