PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 15, 20260 citationsOpen Access

Enhancing Long-Horizon Household Robot Task Planning Using Vision-Language Models

View Full Paper
SBSultan Baig

Key Points

  • To develop and evaluate a reproducible closed-loop embodied AI framework using vision-language models for long-horizon household robot task planning and failure recovery.
  • Built a symbolic household simulator incorporating scene graphs, container affordances, preconditioned manipulation primitives, and a swappable vision-language model interface.
  • Evaluated hierarchical, flat, reactive, and open-loop planning architectures across 160 controlled episodes under simulated stochastic execution noise.
  • Recovery-enabled closed-loop planners achieved a 100% task success rate across all evaluated noise levels.
  • Open-loop planning performance degraded significantly as action failure rates increased, demonstrating an inability to manage stochastic execution errors without dynamic replanning.

Abstract

This paper presents VLM Household Planner, a reproducible embodied AI benchmark and closed-loop agent for long-horizon household robot task planning. The system uses a symbolic household simulator with a scene graph, container affordances, and preconditioned manipulation primitives, together with a swappable Vision-Language Model (VLM) interface. A hierarchical planner decomposes household goals into high-, mid-, and low-level actions, while a closed-loop execution agent detects failures, inserts corrective actions, and retries before continuing the remaining plan.The study evaluates hierarchical, flat, reactive, and open-loop planning approaches across 160 controlled episodes under stochastic execution noise. Recovery-enabled planners achieve 100% task success across the tested noise levels, while open-loop execution degrades substantially as action failure rates increase. The work provides a lightweight, reproducible research and teaching artefact for studying long-horizon planning, task decomposition, multi-step reasoning, and failure recovery in embodied AI and VLM-based household robotics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sultan Baig (2026) studied this question.

synapsesocial.com/papers/6a8019fe75c2e31742c86540https://doi.org/10.5281/zenodo.21923199
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models2025 · 1 citations
  2. 2Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments2024 · 1 citations
  3. 3Planning with Reasoning using Vision Language World Model2025 · 1 citations
  4. 4Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback2024 · 2 citations
  5. 5ReplanVLM: Replanning Robotic Tasks with Visual Language Models2024