Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
June 26, 2023Open Access

Flow to Control: Offline Reinforcement Learning with Lossless Primitive Discovery

View Full Paper
Ask AI
Bookmark
Share

Authors

YYYiqin YangHHHao HuWLWenzhe Li

Discussion

Loading...

Member takes

Overview

Benchmark evaluation demonstrates superior hierarchical control in offline reinforcement learning, indicating that lossless primitives preserve essential policy expressiveness.

Key Points

  • To analyze representation loss in offline hierarchical reinforcement learning and develop a mechanism to discover expressive, lossless action primitives.
  • Formulated a quantitative analysis of policy expressiveness in offline hierarchical reinforcement learning.
  • Implemented a flow-based architecture for low-level policy representations to fully recover the original policy space from static datasets.
  • Evaluated agent performance against existing baselines using the standard D4RL benchmark suite alongside extensive ablation studies.
  • Flow-based primitive representation successfully captured behavioral datasets without restricting coverage of the overall policy space.
  • Achieved superior performance across the majority of evaluated D4RL benchmark tasks compared to prior hierarchical offline methods.

Cite This Study

Yang et al. (2023) studied this question.

synapsesocial.com/papers/6a104f7201be78fe8160b48ehttps://doi.org/10.1609/aaai.v37i9.26286
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Flow-Based Policy for Online Reinforcement Learning2025
  2. 2Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model2026
  3. 3Out of Distribution Adaptation in Offline RL via Causal Normalizing Flows2025
  4. 4Is Value Learning Really the Main Bottleneck in Offline RL?2024
  5. 5Improving the Offline Dataset on Offline Reinforcement Learning2026