PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 24, 20243 citationsOpen Access

SUF: Stabilized Unconstrained Fine-Tuning for Offline-to-Online Reinforcement Learning

View Full Paper
JFJiaheng FengMFMingxiao FengHSHaolin Song

Key Points

Key points are not available for this paper at this time.

Abstract

Offline-to-online reinforcement learning (RL) provides a promising solution to improving suboptimal offline pre-trained policies through online fine-tuning. However, one efficient method, unconstrained fine-tuning, often suffers from severe policy collapse due to excessive distribution shift. To ensure stability, existing methods retain offline constraints and employ additional techniques during fine-tuning, which hurts efficiency. In this work, we introduce a novel perspective: eliminating the policy collapse without imposing constraints. We observe that such policy collapse arises from the mismatch between unconstrained fine-tuning and the conventional RL training framework. To this end, we propose Stabilized Unconstrained Fine-tuning (SUF), a streamlined framework that benefits from the efficiency of unconstrained fine-tuning while ensuring stability by modifying the Update-To-Data ratio. With just a few lines of code adjustments, SUF demonstrates remarkable adaptability to diverse backbones and superior performance over state-of-the-art baselines.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Feng et al. (2024) studied this question.

synapsesocial.com/papers/68e7296db6db6435876a39b5https://doi.org/10.1609/aaai.v38i11.29083
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Conservative Q-Learning for Offline Reinforcement Learning2020 · 538 citations
  2. 2PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning2023 · 5 citations
  3. 3Extreme Q-Learning: MaxEnt RL without Entropy2023 · 5 citations
  4. 4Efficient Deep Reinforcement Learning Requires Regulating Overfitting2023 · 9 citations
  5. 5Dynamic Update-to-Data Ratio: Minimizing World Model Overfitting2023 · 1 citations