PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20250 citationsOpen Access

World-aware Planning Narratives Enhance Large Vision-Language Model Planner

View Full Paper
JSJunhao ShiZFZhaoye FeiSWSiyin Wang

Key Points

  • Task success rates improved by 60.7 with the new planning narrative framework.
  • The framework enhances commonsense reasoning by 60.0 and long-horizon planning by 70.0.
  • Evaluations on the EB-ALFRED benchmark highlight substantial performance gains over existing models.
  • Open-source models outperform proprietary systems, indicating significant advancements in LVLM capabilities.

Abstract

Large Vision-Language Models (LVLMs) show promise for embodied planning tasks but struggle with complex scenarios involving unfamiliar environments and multi-step goals. Current approaches rely on environment-agnostic imitation learning that disconnects instructions from environmental contexts, causing models to struggle with context-sensitive instructions and rely on supplementary cues rather than visual reasoning during long-horizon interactions. In this work, we propose World-Aware Planning Narrative Enhancement (WAP), a framework that infuses LVLMs with comprehensive environmental understanding through four cognitive capabilities (visual appearance modeling, spatial reasoning, functional abstraction, and syntactic grounding) while developing and evaluating models using only raw visual observations through curriculum learning. Evaluations on the EB-ALFRED benchmark demonstrate substantial improvements, with Qwen2.5-VL achieving a 60.7 absolute improvement in task success rates, particularly in commonsense reasoning (+60.0) and long-horizon planning (+70.0). Notably, our enhanced open-source models outperform proprietary systems like GPT-4o and Claude-3.5-Sonnet by a large margin.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Shi et al. (2025) studied this question.

synapsesocial.com/papers/68f04acce559138a1a06e7d0https://doi.org/10.48550/arxiv.2506.21230
Ask AI
Helpful
Bookmark
Share
View Full Paper