PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 28, 20240 citationsOpen Access

Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning

View Full Paper
TZTianle ZhangJGJiayi GuanLZLin Zhao

Key Points

  • The proposed preferred-action-optimized diffusion policy achieves better performance in offline reinforcement learning tasks.
  • Competitive results were demonstrated in sparse reward environments like Kitchen and AntMaze, compared to other methods.
  • Assessed using conditional diffusion models, the approach generates preferred actions, enhancing policy effectiveness during training. The method uses anti-noise preference optimization for stability in learning.

Abstract

Offline reinforcement learning (RL) aims to learn optimal policies from previously collected datasets. Recently, due to their powerful representational capabilities, diffusion models have shown significant potential as policy models for offline RL issues. However, previous offline RL algorithms based on diffusion policies generally adopt weighted regression to improve the policy. This approach optimizes the policy only using the collected actions and is sensitive to Q-values, which limits the potential for further performance enhancement. To this end, we propose a novel preferred-action-optimized diffusion policy for offline RL. In particular, an expressive conditional diffusion model is utilized to represent the diverse distribution of a behavior policy. Meanwhile, based on the diffusion model, preferred actions within the same behavior distribution are automatically generated through the critic function. Moreover, an anti-noise preference optimization is designed to achieve policy improvement by using the preferred actions, which can adapt to noise-preferred actions for stable training. Extensive experiments demonstrate that the proposed method provides competitive or superior performance compared to previous state-of-the-art offline RL methods, particularly in sparse reward tasks such as Kitchen and AntMaze. Additionally, we empirically prove the effectiveness of anti-noise preference optimization.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2024) studied this question.

synapsesocial.com/papers/68e68232b6db64358760ba34https://doi.org/10.48550/arxiv.2405.18729
Ask AI
Helpful
Bookmark
Share
View Full Paper