Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 15, 2025Open Access

Flow-Based Policy for Online Reinforcement Learning

View Full Paper
Ask AI
Bookmark
Share

Authors

LLLei LvYLYunfei LiYLYu Luo

Discussion

Loading...

Member takes

Overview

FlowRL demonstrates competitive reinforcement learning performance, suggesting enhanced action distributions and effective Q-learning.

Key Points

  • FlowRL aligns flow optimization with reinforcement learning objectives, enabling improved policy learning.
  • Empirical evaluations on DMControl and Humanoidbench show FlowRL achieves competitive benchmarks in online reinforcement learning.
  • The framework utilizes a state-dependent velocity field to model policies, generating actions from noise.
  • By bounding the Wasserstein-2 distance, FlowRL maintains proximity to an optimal behavior policy derived from the replay buffer.

Cite This Study

Lv et al. (2025) studied this question.

synapsesocial.com/papers/68efa18f9d05deea71d13cf2https://doi.org/10.48550/arxiv.2506.12811
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Unleashing Flow Policies with Distributional Critics2025
  2. 2One-Step Flow Policy Mirror Descent2025
  3. 3Flow to Control: Offline Reinforcement Learning with Lossless Primitive Discovery2023 · 2 citations
  4. 4GFlowNet Training by Policy Gradients2024
  5. 5Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data2025