PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 12, 20240 citationsOpen Access

VLind-Bench: Measuring Language Priors in Large Vision-Language Models

View Full Paper
KLKang-il LeeMKMinbeom KimSYSeunghyun Yoon

Key Points

Key points are not available for this paper at this time.

Abstract

Large Vision-Language Models (LVLMs) have demonstrated outstanding performance across various multimodal tasks. However, they suffer from a problem known as language prior, where responses are generated based solely on textual patterns while disregarding image information. Addressing the issue of language prior is crucial, as it can lead to undesirable biases or hallucinations when dealing with images that are out of training distribution. Despite its importance, current methods for accurately measuring language priors in LVLMs are poorly studied. Although existing benchmarks based on counterfactual or out-of-distribution images can partially be used to measure language priors, they fail to disentangle language priors from other confounding factors. To this end, we propose a new benchmark called VLind-Bench, which is the first benchmark specifically designed to measure the language priors, or blindness, of LVLMs. It not only includes tests on counterfactual images to assess language priors but also involves a series of tests to evaluate more basic capabilities such as commonsense knowledge, visual perception, and commonsense biases. For each instance in our benchmark, we ensure that all these basic tests are passed before evaluating the language priors, thereby minimizing the influence of other factors on the assessment. The evaluation and analysis of recent LVLMs in our benchmark reveal that almost all models exhibit a significant reliance on language priors, presenting a strong challenge in the field.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lee et al. (2024) studied this question.

synapsesocial.com/papers/68e651cbb6db6435875e2781https://doi.org/10.48550/arxiv.2406.08702
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model2024 · 2 citations
  2. 2VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model2026 · 1 citations
  3. 3Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding2025 · 1 citations
  4. 4DiffuSyn Bench: Evaluating Vision-Language Models on Real-World Complexities with Diffusion-Generated Synthetic Benchmarks2024
  5. 5B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions2024 · 3 citations