PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 9, 20250 citationsOpen Access

Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs

View Full Paper
FZFangrui ZhuHWH. WangYXYiming Xie

Key Points

  • Spatial reasoning abilities are enhanced through structured 2D inputs, indicating effective interaction with 3D environments.
  • In-depth zero-shot analysis demonstrated that closed-source MLLMs can handle route planning and dense captioning tasks.
  • Struct2D utilizes a perception-guided framework to analyze spatial reasoning, relying on 2D representations without 3D data.
  • The findings highlight the potential for bridging perception and language reasoning in large language models, demanding further exploration.

Abstract

Unlocking spatial reasoning in Multimodal Large Language Models (MLLMs) is crucial for enabling intelligent interaction with 3D environments. While prior efforts often rely on explicit 3D inputs or specialized model architectures, we ask: can MLLMs reason about 3D space using only structured 2D representations derived from perception? We introduce Struct2D, a perception-guided prompting framework that combines bird's-eye-view (BEV) images with object marks and object-centric metadata, optionally incorporating egocentric keyframes when needed. Using Struct2D, we conduct an in-depth zero-shot analysis of closed-source MLLMs (e.g., GPT-o3) and find that they exhibit surprisingly strong spatial reasoning abilities when provided with structured 2D inputs, effectively handling tasks such as relative direction estimation and route planning. Building on these insights, we construct Struct2D-Set, a large-scale instruction tuning dataset with 200K fine-grained QA pairs across eight spatial reasoning categories, generated automatically from 3D indoor scenes. We fine-tune an open-source MLLM (Qwen2.5VL) on Struct2D-Set, achieving competitive performance on multiple benchmarks, including 3D question answering, dense captioning, and object grounding. Our approach demonstrates that structured 2D inputs can effectively bridge perception and language reasoning in MLLMs-without requiring explicit 3D representations as input. We will release both our code and dataset to support future research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhu et al. (2025) studied this question.

synapsesocial.com/papers/690fdcdaf60c54d04ea37f50https://doi.org/10.48550/arxiv.2506.04220
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1LLMI3D: Empowering LLM with 3D Perception from a Single 2D Image2024 · 2 citations
  2. 2Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models2025
  3. 3SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models2025
  4. 4How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM2025 · 8 citations
  5. 5Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy2025