PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026Journal of Intelligent & Fuzzy Systems0 citations

Hybrid Soft Actor-Critic with Curriculum Learning for Sparse-Reward Mobile Robot Navigation

View Full Paper
FRFabio Demo RosaRSRaul SteinmetzDGDaniel Fernando Tello Gamarra

Key Points

  • The aim is to improve mobile robot navigation through hybrid reinforcement learning methods by addressing challenges posed by sparse rewards.
  • Evaluated extended Soft Actor-Critic methods for TurtleBot3 navigation in Gazebo.
  • Introduced SAC-XH, which integrates auxiliary shaping signals and a curriculum for better exploration.
  • Conducted experiments in progressively complex Gazebo environments.
  • Implemented a stage-wise Curriculum Learning protocol with competence-based advancement.
  • SAC-XH improved training stability and success rate compared to other methods like SAC, TD3, and DDPG.
  • Achieved success rates between 87-91% under calibrated thresholds with the curriculum protocol.
  • Demonstrated improved learning efficiency and generalization across stages compared to non-curriculum training.

Abstract

This paper presents a unified empirical study of extended Soft Actor-Critic methods for sparse-reward TurtleBot3 navigation in Gazebo under dense 360 ∘ LiDAR observations. We introduce SAC-XH, a streamlined SAC extension that augments the sparse task reward with auxiliary shaping signals and integrates a stage-wise curriculum to improve exploration and sample efficiency. Across progressively complex Gazebo environments, SAC-XH improves training stability and success rate compared to SAC, TD3, and DDPG, while maintaining full reproducibility through an open-source ROS 2/Gazebo framework. SAC-XH consistently outperforms the baselines in learning efficiency and success rate, with dense LiDAR observations (360 beams). Additionally, we evaluate a stage-wise Curriculum Learning protocol on top of SAC-XH, using competence-based advancement and controlled replay transfer. Under calibrated thresholds, the curriculum yields stable convergence and high success rates (87–91%), improving generalization across stages compared to non-curriculum training. These results demonstrate that SAC-XH improves convergence and generalization across multiple Gazebo-simulated navigation environments under sparse-reward conditions, providing a strong DRL baseline for autonomous navigation and a reproducible benchmark for future research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rosa et al. (2026) studied this question.

synapsesocial.com/papers/69be35946e48c4981c673eedhttps://doi.org/10.1177/18758967261431354
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Human-level control through deep reinforcement learning2015 · 31,023 citations
  2. 2Measuring Catastrophic Forgetting in Neural Networks2018 · 574 citations
  3. 3Realistic Counterfactual Explanations for Machine Learning-Controlled Mobile Robots using 2D LiDAR2025 · 1 citations
  4. 4Path Planning of Mobile Robot Based on Improved TD3 Algorithm2022 · 25 citations
  5. 5A deep residual reinforcement learning algorithm based on Soft Actor-Critic for autonomous navigation2024 · 35 citations