PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 2026Robotica1 citations

Adaptive curriculum reinforcement learning with sim-to-real strategy in balance control of underactuated triple pendulum robots

View Full Paper
YFYunfan FuSouthern University of Science and TechnologyJGJing GuoSouthern University of Science and TechnologyDLDonghao LiSouthern University of Science and Technology

Key Points

  • The aim is to develop a model-free RL strategy for balance control in underactuated triple pendulum robots.
  • Utilized a curriculum-based Soft Actor-Critic strategy with an enhanced reward function.
  • Incorporated integral of cumulative joint angle errors in the reward for precision.
  • Implemented motor friction identification and domain randomization for robustness during training.
  • Conducted simulation experiments to assess performance under various challenges.
  • Achieved significant reduction in steady-state errors and improved control precision.
  • Demonstrated superior performance in managing larger initial joint deviations.
  • Successfully maintained balance despite dynamic randomization and sensor noise.
  • Trained policy effectively deployed on a UTPR prototype under real-world conditions.

Abstract

Abstract This paper addresses the challenge of balance control for the underactuated triple pendulum robot (UTPR) using a model-free reinforcement learning (RL) strategy. A curriculum-based Soft Actor-Critic strategy, with a quadratic form and an integral term in the reward function (CSAC-QI), is proposed. By incorporating the integral of cumulative joint angle errors into the reward function, the CSAC-QI method significantly reduces steady-state errors and enhances control precision. CSAC-QI improves convergence efficiency through an adaptive curriculum learning (CL) framework that enables a structured transition from simpler to more complex tasks. To enhance control robustness, motor friction identification and domain randomization are implemented during training, thereby equipping the UTPR to cope with real-world uncertainties. Simulation experiments demonstrate superior performance of the CSAC-QI method in handling larger initial joint deviations, achieving accurate end-effector positioning, and maintaining balance under dynamic randomization, sensor noise, and external disturbances. Notably, the trained policy is directly deployed on the UTPR prototype, where it successfully maintains balance in real-world conditions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fu et al. (2026) studied this question.

synapsesocial.com/papers/69c8c43ede0f0f753b39ee9ehttps://doi.org/10.1017/s0263574726103282
Ask AI
Helpful
Bookmark
Share
View Full Paper