PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Fine-Grained Captioning of Long Videos through Scene Graph Consolidation

View Full Paper
SCSang-Hyeok ChuSSSeonguk SeoBHBohyung Han

Key Points

  • Our method improves video captioning by utilizing graph consolidation for coherent long video summaries.
  • It generates segment-level captions and parses them into scene graphs, achieving significant zero-shot performance.
  • The novel approach reduces computational costs compared to existing LLM-based consolidation methods.
  • By extending the temporal understanding of current models, it alleviates the need for additional fine-tuning.

Abstract

Recent advances in vision-language models have led to impressive progress in caption generation for images and short video clips. However, these models remain constrained by their limited temporal receptive fields, making it difficult to produce coherent and comprehensive captions for long videos. While several methods have been proposed to aggregate information across video segments, they often rely on supervised fine-tuning or incur significant computational overhead. To address these challenges, we introduce a novel framework for long video captioning based on graph consolidation. Our approach first generates segment-level captions, corresponding to individual frames or short video intervals, using off-the-shelf visual captioning models. These captions are then parsed into individual scene graphs, which are subsequently consolidated into a unified graph representation that preserves both holistic context and fine-grained details throughout the video. A lightweight graph-to-text decoder then produces the final video-level caption. This framework effectively extends the temporal understanding capabilities of existing models without requiring any additional fine-tuning on long video datasets. Experimental results show that our method significantly outperforms existing LLM-based consolidation approaches, achieving strong zero-shot performance while substantially reducing computational costs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chu et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcd68d54a28a75cf2211https://doi.org/10.48550/arxiv.2502.16427
Ask AI
Helpful
Bookmark
Share
View Full Paper