PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 10, 20250 citationsOpen Access

Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation

View Full Paper
CWCongcong WenGBGeeta Chandra Raju BethalaHYHao Yu

Key Points

  • The humanoid-COA framework enables effective zero-shot loco-manipulation in robots, marking a significant advancement in robotic assistance.
  • Extensive experimentation reveals that humanoid-COA outperforms existing models in locomotion and manipulation tasks by a substantial margin.
  • Reasoning in humanoid-COA transforms high-level human instructions into actionable locomotion and manipulation sequences, enhancing robot adaptability.
  • Effective deployment of the humanoid agent relies on its ability to generalize across various tasks and environments, solving a critical robotic challenge.

Abstract

Humanoid loco-manipulation, which integrates whole-body locomotion with dexterous manipulation, remains a fundamental challenge in robotics. Beyond whole-body coordination and balance, a central difficulty lies in understanding human instructions and translating them into coherent sequences of embodied actions. Recent advances in foundation models provide transferable multimodal representations and reasoning capabilities, yet existing efforts remain largely restricted to either locomotion or manipulation in isolation, with limited applicability to humanoid settings. In this paper, we propose Humanoid-COA, the first humanoid agent framework that integrates foundation model reasoning with an Embodied Chain-of-Action (CoA) mechanism for zero-shot loco-manipulation. Within the perception--reasoning--action paradigm, our key contribution lies in the reasoning stage, where the proposed CoA mechanism decomposes high-level human instructions into structured sequences of locomotion and manipulation primitives through affordance analysis, spatial inference, and whole-body action reasoning. Extensive experiments on two humanoid robots, Unitree H1-2 and G1, in both an open test area and an apartment environment, demonstrate that our framework substantially outperforms prior baselines across manipulation, locomotion, and loco-manipulation tasks, achieving robust generalization to long-horizon and unstructured scenarios. Project page: https://humanoid-coa.github.io/

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wen et al. (2025) studied this question.

synapsesocial.com/papers/68e861a57ef2f04ca37e4586https://doi.org/10.48550/arxiv.2504.09532
Ask AI
Helpful
Bookmark
Share
View Full Paper