PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 12, 20250 citationsOpen Access

Learning to Walk with Less: a Dyna-Style Approach to Quadrupedal Locomotion

View Full Paper
FPF. A. C. PintoFTFelipe Andrade G. TommaselliJNJuliano Decico Negri

Key Points

  • Our approach enhances sample efficiency through synthetic data integration in quadrupedal locomotion.
  • A predictive model correlated rollout length with efficiency, enabling targeted experiments and improved policy returns.
  • Utilizing synthetic transitions while maintaining fidelity, we reduced variance in locomotion performance metrics.
  • Successful validation on the Unitree Go1 robot showcases practical improvements in adaptive tracking capabilities.

Abstract

Traditional RL-based locomotion controllers often suffer from low data efficiency, requiring extensive interaction to achieve robust performance. We present a model-based reinforcement learning (MBRL) framework that improves sample efficiency for quadrupedal locomotion by appending synthetic data to the end of standard rollouts in PPO-based controllers, following the Dyna-Style paradigm. A predictive model, trained alongside the policy, generates short-horizon synthetic transitions that are gradually integrated using a scheduling strategy based on the policy update iterations. Through an ablation study, we identified a strong correlation between sample efficiency and rollout length, which guided the design of our experiments. We validated our approach in simulation on the Unitree Go1 robot and showed that replacing part of the simulated steps with synthetic ones not only mimics extended rollouts but also improves policy return and reduces variance. Finally, we demonstrate that this improvement transfers to the ability to track a wide range of locomotion commands using fewer simulated steps.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pinto et al. (2025) studied this question.

synapsesocial.com/papers/68ec1be02b8fa9b2b78ad00ehttps://doi.org/10.48550/arxiv.2509.06296
Ask AI
Helpful
Bookmark
Share
View Full Paper