PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 14, 20240 citationsOpen Access

NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

View Full Paper
PPPranshu PandyaATAgney S TalwarrVGVatsal Gupta

Key Points

Key points are not available for this paper at this time.

Abstract

Cognitive textual and visual reasoning tasks, such as puzzles, series, and analogies, demand the ability to quickly reason, decipher, and evaluate patterns both textually and spatially. While LLMs and VLMs, through extensive training on large amounts of human-curated data, have attained a high level of pseudo-human intelligence in some common sense reasoning tasks, they still struggle with more complex reasoning tasks that require cognitive understanding. In this work, we introduce a new dataset, NTSEBench, designed to evaluate the cognitive multi-modal reasoning and problem-solving skills of large models. The dataset comprises 2,728 multiple-choice questions comprising of a total of 4,642 images across 26 categories sampled from the NTSE examination conducted nationwide in India, featuring both visual and textual general aptitude questions that do not rely on rote learning. We establish baselines on the dataset using state-of-the-art LLMs and VLMs. To facilitate a comparison between open source and propriety models, we propose four distinct modeling strategies to handle different modalities (text and images) in the dataset instances.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pandya et al. (2024) studied this question.

synapsesocial.com/papers/68e60662b6db643587599d9chttps://doi.org/10.48550/arxiv.2407.10380
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models2024
  2. 2ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models2025
  3. 3ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models2025
  4. 4TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning2025
  5. 5Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models2024 · 3 citations