PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 24, 20243 citationsOpen Access

SUF: Stabilized Unconstrained Fine-Tuning for Offline-to-Online Reinforcement Learning

View Full Paper
JFJiaheng FengMFMingxiao FengHSHaolin Song

Key Points

Key points are not available for this paper at this time.

Abstract

Offline-to-online reinforcement learning (RL) provides a promising solution to improving suboptimal offline pre-trained policies through online fine-tuning. However, one efficient method, unconstrained fine-tuning, often suffers from severe policy collapse due to excessive distribution shift. To ensure stability, existing methods retain offline constraints and employ additional techniques during fine-tuning, which hurts efficiency. In this work, we introduce a novel perspective: eliminating the policy collapse without imposing constraints. We observe that such policy collapse arises from the mismatch between unconstrained fine-tuning and the conventional RL training framework. To this end, we propose Stabilized Unconstrained Fine-tuning (SUF), a streamlined framework that benefits from the efficiency of unconstrained fine-tuning while ensuring stability by modifying the Update-To-Data ratio. With just a few lines of code adjustments, SUF demonstrates remarkable adaptability to diverse backbones and superior performance over state-of-the-art baselines.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Feng et al. (2024) studied this question.

synapsesocial.com/papers/68e7296db6db6435876a39b5https://doi.org/10.1609/aaai.v38i11.29083
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Offline-to-Online Reinforcement Learning via Balanced Replay and\n Pessimistic Q-Ensemble2021 · 18 citations
  2. 2Actor-Critic Alignment for Offline-to-Online Reinforcement Learning.2023 · 2 citations
  3. 3PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning2023 · 5 citations
  4. 4Extreme Q-Learning: MaxEnt RL without Entropy2023 · 5 citations
  5. 5Improving TD3-BC: Relaxed Policy Constraint for Offline Learning and Stable Online Fine-Tuning2022 · 8 citations