Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 1, 2025Open Access

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs

View Full Paper
Ask AI
Bookmark
Share

Authors

TMTianwei MaYZYuanxing ZhangJRJincheng Ren

Discussion

Loading...

Member takes

Overview

IV-Bench evaluates image-grounded video reasoning in MLLMs, highlighting significant performance gaps.

Key Points

  • Current models achieve only up to 28.9% accuracy on video perception and reasoning tasks.
  • IV-Bench consists of 967 videos and 2,585 annotated image-text queries covering 13 distinct tasks.
  • Key performance factors include inference patterns, frame numbers, and resolution in videos.
  • Findings underscore the need for improved approaches in aligning data formats and refining model training.

Cite This Study

Ma et al. (2025) studied this question.

synapsesocial.com/papers/68dd91c7fe798ba2fc498446https://doi.org/10.48550/arxiv.2504.15415
View Full Paper
Ask AI
Bookmark
Share