PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 30, 20241 citationsOpen Access

Diffusion Policies creating a Trust Region for Offline Reinforcement Learning

View Full Paper
TCTianyu ChenZWZhendong WangMZMingyuan Zhou

Key Points

  • DTQL leads to faster training and inference without sacrificing performance, achieving superior results in various benchmark tasks.
  • Key improvements include eliminating the need for iterative denoising sampling, which enhances computational efficiency significantly.
  • Assessment using multiple 2D bandit scenarios indicates DTQL's robust performance compared to Kullback-Leibler distillation strategies overall in testing environments across multiple gym tasks.  The introduction of a trust region loss crucially links the diffusion policy and the one-step policy, promoting effective exploration.

Abstract

Offline reinforcement learning (RL) leverages pre-collected datasets to train optimal policies. Diffusion Q-Learning (DQL), introducing diffusion models as a powerful and expressive policy class, significantly boosts the performance of offline RL. However, its reliance on iterative denoising sampling to generate actions slows down both training and inference. While several recent attempts have tried to accelerate diffusion-QL, the improvement in training and/or inference speed often results in degraded performance. In this paper, we introduce a dual policy approach, Diffusion Trusted Q-Learning (DTQL), which comprises a diffusion policy for pure behavior cloning and a practical one-step policy. We bridge the two polices by a newly introduced diffusion trust region loss. The diffusion policy maintains expressiveness, while the trust region loss directs the one-step policy to explore freely and seek modes within the region defined by the diffusion policy. DTQL eliminates the need for iterative denoising sampling during both training and inference, making it remarkably computationally efficient. We evaluate its effectiveness and algorithmic characteristics against popular Kullback-Leibler (KL) based distillation methods in 2D bandit scenarios and gym tasks. We then show that DTQL could not only outperform other methods on the majority of the D4RL benchmark tasks but also demonstrate efficiency in training and inference speeds. The PyTorch implementation will be made available.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2024) studied this question.

synapsesocial.com/papers/68e67aa1b6db643587604f01https://doi.org/10.48550/arxiv.2405.19690
Ask AI
Helpful
Bookmark
Share
View Full Paper