Framework enhances decision-making in web agents, indicating improvements in complex web navigation tasks.
Large language model (LLM)-based agents have shown remarkable potential in automating webbased tasks, but struggle significantly with complex, long-horizon web navigation and interaction scenarios characterized by vast action spaces, dynamic changes, and multi-step task dependencies. To address these challenges, we introduce HiPlan, a hierarchical planning framework that provides adaptive global-local guidance to boost web agents’ decision-making. HiPlan first decomposes high-level objectives into milestone action guides for maintaining navigation direction and progress, while simultaneously offering step-wise hints for fine-grained and real-time interaction decisions that effectively bridge gaps and correct navigation errors. To enhance generalization and robustness across diverse tasks, we propose a dual-level experience reuse mechanism that leverages expert demonstrations to guide both the decomposition of the task and the generation of hints aligned with each milestone objective. This mechanism enables the agent to learn from successful historical interactions while adapting to current task contexts. Extensive experiments on challenging web agent benchmarks demonstrate that HiPlan substantially outperforms strong baselines, with ablation studies validating the critical and complementary benefits of all components for web-based agent planning.
No takes yet. Share an insight, caveat, or question.
Li et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: