PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

Online Knowledge Distillation with Reward Guidance

View Full Paper
JCJia Chen

Key Points

  • The proposed framework significantly reduces the performance gap between student and teacher policies.
  • Min-max optimization is utilized within the reward-guided imitation learning approach for effective knowledge distillation.
  • Theoretical analysis and empirical results support the framework's efficacy in preference-based knowledge distillation.
  • Moreover, enhancements to the reward model extend its application to white-box knowledge distillation setups.

Abstract

This work studies knowledge distillation (KD) for large language models (LLMs) through preference optimization. We propose a reward-guided imitation learning framework for sequential KD, formulating a min-max optimization problem between the policy and reward model (RM) to minimize the performance gap between the student and teacher policies. Specifically, the reward optimization is constrained to achieve near-optimality within a confidence set for preference alignment. For preference data construction, we explore both offline and online preference-based KD. Additionally, we reformulate the RM using the Q-value function and extend the framework to white-box KD, where the teacher policy's predicted probabilities are accessible. Theoretical analysis and empirical results demonstrate the effectiveness of the proposed framework.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jia Chen (2025) studied this question.

synapsesocial.com/papers/68da58d8c1728099cfd1126chttps://doi.org/10.48550/arxiv.2505.18952
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Direct Preference Knowledge Distillation for Large Language Models2024
  2. 2Adversarial Moment-Matching Distillation of Large Language Models2024
  3. 3Revisiting Knowledge Distillation for Autoregressive Language Models2024
  4. 4PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs2024
  5. 5Knowledge Editing in Language Models via Adapted Direct Preference Optimization2024