PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 20, 20241 citationsOpen Access

Bayesian Reward Models for LLM Alignment

View Full Paper
AYAdam X. YangMRMaxime RobeynsTCThomas Coste

Key Points

Key points are not available for this paper at this time.

Abstract

To ensure that large language model (LLM) responses are helpful and non-toxic, we usually fine-tune a reward model on human preference data. We then select policy responses with high rewards (best-of-n sampling) or further optimize the policy to produce responses with high rewards (reinforcement learning from human feedback). However, this process is vulnerable to reward overoptimization or hacking, in which the responses selected have high rewards due to errors in the reward model rather than a genuine preference. This is especially problematic as the prompt or response diverges from the training data. It should be possible to mitigate these issues by training a Bayesian reward model, which signals higher uncertainty further from the training data distribution. Therefore, we trained Bayesian reward models using Laplace-LoRA (Yang et al., 2024) and found that the resulting uncertainty estimates can successfully mitigate reward overoptimization in best-of-n sampling.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2024) studied this question.

synapsesocial.com/papers/68e786ffb6db6435876f9c2ahttps://doi.org/10.48550/arxiv.2402.13210
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Fine-Tuning Language Models with Reward Learning on Policy2024 · 1 citations
  2. 2Towards Reliable, Uncertainty-Aware Alignment2025
  3. 3Robust Reinforcement Learning from Human Feedback for Large Language Models Fine-Tuning2025
  4. 4Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs2024
  5. 5Aligning Crowd Feedback via Distributional Preference Reward Modeling2024 · 1 citations