PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20250 citationsOpen Access

X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real

View Full Paper
PDPrithwish DanCornell UniversityKKKushal KediaIndian Institute of Technology KharagpurACAngela ChaoCornell University

Key Points

  • X-Sim improves task progress by 30% over traditional hand-tracking methods and sim-to-real baselines.
  • The framework matches behavior cloning performance with 10x less data collection time, making it highly efficient.
  • X-Sim generalizes effectively to new camera viewpoints and different test-time conditions, enhancing its adaptability.
  • No robot teleoperation data is required, simplifying the learning process for complex tasks.

Abstract

Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment approaches try to map human motion to robot actions, but often fail when the embodiments differ significantly. We propose X-Sim, a real-to-sim-to-real framework that uses object motion as a dense and transferable signal for learning robot policies. X-Sim starts by reconstructing a photorealistic simulation from an RGBD human video and tracking object trajectories to define object-centric rewards. These rewards are used to train a reinforcement learning (RL) policy in simulation. The learned policy is then distilled into an image-conditioned diffusion policy using synthetic rollouts rendered with varied viewpoints and lighting. To transfer to the real world, X-Sim introduces an online domain adaptation technique that aligns real and simulated observations during deployment. Importantly, X-Sim does not require any robot teleoperation data. We evaluate it across 5 manipulation tasks in 2 environments and show that it: (1) improves task progress by 30% on average over hand-tracking and sim-to-real baselines, (2) matches behavior cloning with 10x less data collection time, and (3) generalizes to new camera viewpoints and test-time changes. Code and videos are available at https://portal-cornell.github.io/X-Sim/.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dan et al. (2025) studied this question.

synapsesocial.com/papers/68f147cc724575985c3fd092https://doi.org/10.48550/arxiv.2505.07096
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation2024 · 5 citations
  2. 2TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction2024 · 1 citations
  3. 3EAGERx: Graph-Based Framework for Sim2real Robot Learning2024
  4. 4ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis2025
  5. 5Sim2Real Manipulation on Unknown Objects with Tactile-based Reinforcement Learning2024