PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 31, 20240 citationsOpen Access

Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation

View Full Paper
SGShangding GuLSLaixi ShiYDYuhao Ding

Key Points

Key points are not available for this paper at this time.

Abstract

Safe reinforcement learning (RL) is crucial for deploying RL agents in real-world applications, as it aims to maximize long-term rewards while satisfying safety constraints. However, safe RL often suffers from sample inefficiency, requiring extensive interactions with the environment to learn a safe policy. We propose Efficient Safe Policy Optimization (ESPO), a novel approach that enhances the efficiency of safe RL through sample manipulation. ESPO employs an optimization framework with three modes: maximizing rewards, minimizing costs, and balancing the trade-off between the two. By dynamically adjusting the sampling process based on the observed conflict between reward and safety gradients, ESPO theoretically guarantees convergence, optimization stability, and improved sample complexity bounds. Experiments on the Safety-MuJoCo and Omnisafe benchmarks demonstrate that ESPO significantly outperforms existing primal-based and primal-dual-based baselines in terms of reward maximization and constraint satisfaction. Moreover, ESPO achieves substantial gains in sample efficiency, requiring 25--29% fewer samples than baselines, and reduces training time by 21--38%.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gu et al. (2024) studied this question.

synapsesocial.com/papers/68e6785bb6db643587602a37https://doi.org/10.48550/arxiv.2405.20860
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation2024 · 12 citations
  2. 2Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation2024
  3. 3Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization2024 · 6 citations
  4. 4Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization2024
  5. 5Mirror Descent Safe Policy Optimization for Reinforcement Learning Agents2026 · 1 citations