Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
October 1, 2025IEEE Transactions on Pattern Analysis and Machine Intelligence

A Survey on Video Temporal Grounding with Multimodal Large Language Model

View Full Paper
Ask AI
Bookmark
Share

Authors

JWJianlong WuWLWei LiuYLYe Liu

Discussion

Loading...

Member takes

Overview

This survey examines video temporal grounding methods using multimodal large language models, highlighting their training paradigms and effectiveness.

Key Points

  • VTG-MLLMs outperform traditional methods in competitive performance and generalization across various scenarios.
  • The survey identifies three key aspects: functional roles of MLLMs, training paradigms, and video feature processing techniques.
  • Benchmark datasets and evaluation protocols are discussed, along with empirical findings in the field.
  • Future directions for research are proposed, detailing limitations in current approaches to VTG.

Cite This Study

Wu et al. (2025) studied this question.

synapsesocial.com/papers/68dd7e78fe798ba2fc496139https://doi.org/10.1109/tpami.2025.3615586
View Full Paper
Ask AI
Bookmark
Share