Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 13, 2025Open Access

GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration

View Full Paper
Ask AI
Bookmark
Share

Authors

XWXu WanFZFeng ZhuYZYihan Zeng

Discussion

Loading...

Member takes

Overview

Proposed GLaVE-Cap improves video captioning accuracy by integrating local and global contexts, suggesting stronger local-global interactions.

Key Points

  • GLaVE-Cap improves video captioning by addressing challenges in local-to-global paradigms, leading to more detailed captions.
  • The model incorporates TrackFusion, generating comprehensive local captions while using vision experts for enhanced visual prompts.
  • CaptionBridge facilitates interaction between local and global captions, summarizing them into coherent global descriptions.
  • GLaVE-1.2M introduces a large dataset of fine-grained video captions and related question-answer pairs for more reliable evaluation.

Cite This Study

Wan et al. (2025) studied this question.

synapsesocial.com/papers/68ecfebf950606aabec09469https://doi.org/10.48550/arxiv.2509.11360
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Controllable Hybrid Captioner for Improved Long-form Video Understanding2025
  2. 2Ensemble-based multilingual video captioning with multimodel fusion of visual and audio cues2026
  3. 3FlexCap: Generating Rich, Localized, and Flexible Captions in Images2024
  4. 4EvCap: Element-Aware Video Captioning2024 · 16 citations
  5. 5Video Enriched Retrieval Augmented Generation Using Aligned Video Captions2024