PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 12, 20248 citationsOpen Access

Reference-free Monolithic Preference Optimization with Odds Ratio

View Full Paper
JHJiwoo HongPohang University of Science and TechnologyNLNoah LeeKorea Advanced Institute of Science and TechnologyJTJames H. ThorneKorea Advanced Institute of Science and Technology

Key Points

Key points are not available for this paper at this time.

Abstract

While recent preference alignment algorithms for language models have demonstrated promising results, supervised fine-tuning (SFT) remains imperative for achieving successful convergence. In this paper, we study the crucial role of SFT within the context of preference alignment, emphasizing that a minor penalty for the disfavored generation style is sufficient for preference-aligned SFT. Building on this foundation, we introduce a straightforward and innovative reference model-free monolithic odds ratio preference optimization algorithm, ORPO, eliminating the necessity for an additional preference alignment phase. We demonstrate, both empirically and theoretically, that the odds ratio is a sensible choice for contrasting favored and disfavored styles during SFT across the diverse sizes from 125M to 7B. Specifically, fine-tuning Phi-2 (2. 7B), Llama-2 (7B), and Mistral (7B) with ORPO on the UltraFeedback alone surpasses the performance of state-of-the-art language models with more than 7B and 13B parameters: achieving up to 12. 20% on AlpacaEval₂. ₀ and 7. 32 in MT-Bench, as shown in Figures 1 and 12. We release code and model checkpoints for Mistral-ORPO- (7B) and Mistral-ORPO- (7B).

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hong et al. (2024) studied this question.

synapsesocial.com/papers/68e745afb6db6435876bef33https://doi.org/10.48550/arxiv.2403.07691
Ask AI
Helpful
Bookmark
Share
View Full Paper