Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 10, 2025

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

View Full Paper
Ask AI
Bookmark
Share

Authors

CLChaoyu LiArizona State UniversityEIEun Woo ImArizona State UniversityPFPooyan FazliArizona State University

Discussion

Loading...

Member takes

Implication

Benchmarking hallucinations in multimodal large language models for video understanding, indicating potential improvement strategies.

Key Points

  • Most multimodal large language models experience hallucinations, particularly regarding action and scene transitions.
  • VIDHALLUC, a benchmark with 5,002 videos, reveals critical dimensions where hallucination occurs in model outputs.
  • Testing of DINO-HEAL shows a significant average performance improvement of 3.02% in mitigating hallucinations.
  • DINO-HEAL utilizes spatial saliency to adjust visual features and effectively reduce hallucinations during inference.

Cite This Study

Li et al. (2025) studied this question.

synapsesocial.com/papers/68c1c23554b1d3bfb60efb92https://doi.org/10.1109/cvpr52734.2025.01281
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models2024 · 4 citations
  2. 2Detecting and Preventing Hallucinations in Large Vision Language Models2024 · 146 citations
  3. 3Hallucination of Multimodal Large Language Models: A Survey2024 · 30 citations
  4. 4Visual Hallucinations of Multi-modal Large Language Models2024
  5. 5Detecting and Evaluating Medical Hallucinations in Large Vision Language Models2024 · 7 citations