PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 30, 20240 citationsOpen Access

Learning from Random Demonstrations: Offline Reinforcement Learning with Importance-Sampled Diffusion Models

View Full Paper
ZFZeyu FangTLTian Lan

Key Points

  • The proposed method achieves significant improvement in learning performance compared to existing baselines.
  • It leverages random demonstrations and guided diffusion for effective well-aligned policy evaluation.
  • Evaluation in the D4RL environment shows a notable reduction in return gap with an optimal policy under varying input conditions, emphasizing flexibility in data use and model adaptation. Requires more innovations in world model adaptations for advanced policy reinforcement learning applications. Supports better real-world application by improving learning from random demonstrations.

Abstract

Generative models such as diffusion have been employed as world models in offline reinforcement learning to generate synthetic data for more effective learning. Existing work either generates diffusion models one-time prior to training or requires additional interaction data to update it. In this paper, we propose a novel approach for offline reinforcement learning with closed-loop policy evaluation and world-model adaptation. It iteratively leverages a guided diffusion world model to directly evaluate the offline target policy with actions drawn from it, and then performs an importance-sampled world model update to adaptively align the world model with the updated policy. We analyzed the performance of the proposed method and provided an upper bound on the return gap between our method and the real environment under an optimal policy. The result sheds light on various factors affecting learning performance. Evaluations in the D4RL environment show significant improvement over state-of-the-art baselines, especially when only random or medium-expertise demonstrations are available -- thus requiring improved alignment between the world model and offline policy evaluation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fang et al. (2024) studied this question.

synapsesocial.com/papers/68e67bb1b6db643587605fc4https://doi.org/10.48550/arxiv.2405.19878
Ask AI
Helpful
Bookmark
Share
View Full Paper