PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 30, 2026IEEE Transactions on Pattern Analysis and Machine Intelligence0 citations

YOTO++: Learning Long-Horizon Closed-Loop Bimanual Manipulation from One-Shot Human Video Demonstrations

View Full Paper
HZHuayi ZhouRWRuixiang WangYTYunxin Tai

Key Points

  • This research aims to enhance bimanual robotic manipulation by developing a one-shot learning framework that utilizes human video demonstrations.
  • Developed YOTO++ framework for one-shot learning from third-person video demonstrations.
  • Extracted structured 3D hand motions and created keyframe-based trajectories for dual-arm execution.
  • Implemented a visual alignment mechanism for closed-loop control during manipulation.
  • YOTO++ achieved strong generalization across diverse bimanual tasks, including both synchronized and contact-rich scenarios.
  • The system demonstrated seamless cross-embodiment transfer on an unseen dual-arm robotic platform without retraining.
  • Impressive performance metrics in accuracy and scalability were reported, advancing practical bimanual manipulation.

Abstract

Bimanual robotic manipulation remains a fundamental challenge due to the inherent complexity of dual-arm coordination and high-dimensional action spaces. This paper presents the extended YOTO++ (You Only Teach Once), which is a unified one-shot learning framework for teaching bimanual skills directly from third-person human video demonstrations. Our method extracts structured 3D hand motions using binocular vision and distills them into compact, keyframe-based trajectories for dual-arm execution. We develop a scalable demonstration proliferation strategy that synthetically augments one-shot demonstrations into diverse training samples, enabling effective learning of a customized bimanual diffusion policy. Extensive evaluations across a broad spectrum of long-horizon bimanual tasks, including asynchronous, synchronous, contact-rich, and non-prehensile scenarios, demonstrate strong generalization to novel skills and objects. We further introduce a visual alignment mechanism at the initial manipulation stage for closed-loop control, enabling the system to timely adapt to perturbations during execution. We validate the framework on an unseen dual-arm robotic platform to show seamless cross-embodiment transfer without additional retraining. YOTO++ achieves impressive performance in accuracy, robustness, and scalability, advancing the practical deployment of general-purpose bimanual manipulation systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhou et al. (2026) studied this question.

synapsesocial.com/papers/69f2f0991e5f7920c6386ce2https://doi.org/10.1109/tpami.2026.3688078
Ask AI
Helpful
Bookmark
Share
View Full Paper