PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20251 citationsOpen Access

Spatial Mental Modeling from Limited Views

View Full Paper
BYBangjie YinQWQineng WangPZPingyue Zhang

Key Points

  • Significant improvement in VLM accuracy from 37.8% to 60.8% after implementing spatial mental models and reasoning techniques.
  • The 'map-then-reason' approach effectively combines cognitive mapping and reasoning, leading to an accuracy boost of 23.0%.
  • Using reinforcement learning further enhanced VLM performance to 70.7%, indicating the method's robustness.
  • MindCube benchmark reveals critical gaps in VLM abilities, highlighting the importance of spatial reasoning in AI models.

Abstract

Can Vision Language Models (VLMs) imagine the full scene from just a few views, like humans do? Humans form spatial mental models, internal representations of unseen space, to reason about layout, perspective, and motion. Our new MindCube benchmark with 21,154 questions across 3,268 images exposes this critical gap, where existing VLMs exhibit near-random performance. Using MindCube, we systematically evaluate how well VLMs build robust spatial mental models through representing positions (cognitive mapping), orientations (perspective-taking), and dynamics (mental simulation for "what-if" movements). We then explore three approaches to help VLMs approximate spatial mental models, including unseen intermediate views, natural language reasoning chains, and cognitive maps. The significant improvement comes from a synergistic approach, "map-then-reason", that jointly trains the model to first generate a cognitive map and then reason upon it. By training models to reason over these internal maps, we boosted accuracy from 37.8% to 60.8% (+23.0%). Adding reinforcement learning pushed performance even further to 70.7% (+32.9%). Our key insight is that such scaffolding of spatial mental models, actively constructing and utilizing internal structured spatial representations with flexible reasoning processes, significantly improves understanding of unobservable space.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yin et al. (2025) studied this question.

synapsesocial.com/papers/68f04acce559138a1a06e98chttps://doi.org/10.48550/arxiv.2506.21458
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models2024 · 3 citations
  2. 2Spatial intelligence in vision-language models: a comprehensive survey2026
  3. 3ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models2026 · 2 citations
  4. 4Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models2024 · 2 citations
  5. 5VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs2024