This paper presents VLM Household Planner, a reproducible embodied AI benchmark and closed-loop agent for long-horizon household robot task planning. The system uses a symbolic household simulator with a scene graph, container affordances, and preconditioned manipulation primitives, together with a swappable Vision-Language Model (VLM) interface. A hierarchical planner decomposes household goals into high-, mid-, and low-level actions, while a closed-loop execution agent detects failures, inserts corrective actions, and retries before continuing the remaining plan.The study evaluates hierarchical, flat, reactive, and open-loop planning approaches across 160 controlled episodes under stochastic execution noise. Recovery-enabled planners achieve 100% task success across the tested noise levels, while open-loop execution degrades substantially as action failure rates increase. The work provides a lightweight, reproducible research and teaching artefact for studying long-horizon planning, task decomposition, multi-step reasoning, and failure recovery in embodied AI and VLM-based household robotics.
Sultan Baig (2026) studied this question.