PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

Diffusion Self-Weighted Guidance for Offline Reinforcement Learning

View Full Paper
ATA. Sanchez de TagleJRJavier Ruiz‐del‐SolarFTFelipe Tobar

Key Points

  • Self-Weighted Guidance achieves optimal policy generation in offline reinforcement learning by using a diffusion model.
  • The proposed method aligns guidance functions with the diffusion model, simplifying the computation of required scores.
  • Performance of Self-Weighted Guidance on challenging D4RL environments matches that of leading methods.
  • Ablation studies confirm the scalability and effectiveness of different weight formulations in guiding the learning process.

Abstract

Offline reinforcement learning (RL) recovers the optimal policy given historical observations of an agent. In practice, is modeled as a weighted version of the agent's behavior policy, using a weight function w working as a critic of the agent's behavior. Though recent approaches to offline RL based on diffusion models have exhibited promising results, the computation of the required scores is challenging due to their dependence on the unknown w. In this work, we alleviate this issue by constructing a diffusion over both the actions and the weights. With the proposed setting, the required scores are directly obtained from the diffusion model without learning extra networks. Our main conceptual contribution is a novel guidance method, where guidance (which is a function of w) comes from the same diffusion model, therefore, our proposal is termed Self-Weighted Guidance (SWG). We show that SWG generates samples from the desired distribution on toy examples and performs on par with state-of-the-art methods on D4RL's challenging environments, while maintaining a streamlined training pipeline. We further validate SWG through ablation studies on weight formulations and scalability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tagle et al. (2025) studied this question.

synapsesocial.com/papers/68da58d8c1728099cfd11042https://doi.org/10.48550/arxiv.2505.18345
Ask AI
Helpful
Bookmark
Share
View Full Paper