PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

View Full Paper
RYRuilin YaoFJFrank JiangJHJirui Huang

Key Points

  • Multimodal large language models struggle with reasoning, scoring less than 60% accuracy on complex tasks.
  • Lens benchmark includes 3.4K images and 60K+ questions across eight tasks, focusing on perception and understanding.
  • Evaluation of 15+ frontier MLLMs, including GPT-4o and QVQ-72B-preview, highlights gaps in their reasoning capabilities.
  • Dataset derived from social media images aims to improve multimodal reasoning assessments, calling for more robust methodologies.

Abstract

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. The existing benchmarks are usually constructed in the task-oriented manner without guarantee that different task samples come from the same data distribution, thus they often fall short in evaluating the synergistic effects of lower-level perceptual capabilities on higher-order reasoning. To lift this limitation, we contribute Lens, a multi-level benchmark with 3.4K contemporary images and 60K+ human-authored questions covering eight tasks and 12 daily scenarios, forming three progressive task tiers, i.e., perception, understanding, and reasoning. One feature is that each image is equipped with rich annotations for all tasks. Thus, this dataset intrinsically supports to evaluate MLLMs to handle image-invariable prompts, from basic perception to compositional reasoning. In addition, our images are manully collected from the social media, in which 53% were published later than Jan. 2025. We evaluate 15+ frontier MLLMs such as Qwen2.5-VL-72B, InternVL3-78B, GPT-4o and two reasoning models QVQ-72B-preview and Kimi-VL. These models are released later than Dec. 2024, and none of them achieve an accuracy greater than 60% in the reasoning tasks. Project page: https://github.com/Lens4MLLMs/lens. ICCV 2025 workshop page: https://lens4mllms.github.io/mars2-workshop-iccv2025/

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yao et al. (2025) studied this question.

synapsesocial.com/papers/68f5c338e2d8b12842645b77https://doi.org/10.48550/arxiv.2505.15616
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning2024 · 2 citations
  2. 2NPHardEval4V: A Dynamic Reasoning Benchmark of Multimodal Large Language Models2024 · 1 citations
  3. 3IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs2025
  4. 4A Benchmark for Multi-modal Foundation Models on Low-level Vision: from Single Images to Pairs2024
  5. 5Can MLLMs Guide Me Home? A Benchmark Study on Fine-Grained Visual Reasoning from Transit Maps2025