PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 26, 2025Transactions of the Institute of Measurement and Control0 citationsOpen Access

Research on multi-terrain locomotion control of quadruped robots based on an improved PPO algorithm

View Full Paper
GZG. ZhangNCNaijian ChenYJYiming Ji

Key Points

  • The proposed LSTM-PPO algorithm promotes adaptability and stability for quadruped locomotion on rough terrain.
  • Training in simulation with Isaac Gym led to improved performance metrics, showing faster convergence and higher rewards.
  • A distributed reinforcement learning framework was developed, integrating a multi-objective reward and penalty system.
  • The policy's real-world effectiveness was tested on a quadruped robot, demonstrating robust performance on various complex terrains.

Abstract

To address the challenges of poor body stability and limited generalization in quadruped locomotion over unstructured terrains, this paper proposes a locomotion control method based on an improved Long Short-Term Memory–Proximal Policy Optimization (LSTM-PPO) algorithm. An LSTM-based state processing module is integrated into the policy and value networks to handle the variation in input state length caused by complex terrain transitions. The PPO architecture is accordingly modified to support sequential state encoding. A distributed reinforcement learning framework is constructed, incorporating a multi-objective reward and penalty mechanism to enhance the adaptability of the learned policy across diverse environments. The training is conducted in a simulation environment built on Isaac Gym, where ablation studies are performed to validate the effectiveness of the LSTM module and the rationality of its hidden layer configuration. Comparative experiments against standard PPO, TD3, and DDPG demonstrate that the proposed algorithm achieves faster convergence, higher cumulative rewards, and more stable training. Finally, the learned policy is deployed and validated in both the Gazebo simulation platform and a real quadruped robot. Experimental results show that the proposed method enables the robot to effectively adapt to complex terrains such as grass, slopes, and stairs, exhibiting strong robustness and practical applicability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68d6cd68b1249cec298b3b20https://doi.org/10.1177/01423312251371751
Ask AI
Helpful
Bookmark
Share
View Full Paper