PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
Synapse
⌘+K
Synapse
February 22, 2024Open Access

Generalizing Reward Modeling for Out-of-Distribution Preference Learning

View Full Paper
Ask AI
Bookmark
Share

Authors

JCJia Chen

Discussion

Loading...

Member takes

Overview

Key Points

Key points are not available for this paper at this time.

Cite This Study

Jia Chen (2024) studied this question.

synapsesocial.com/papers/68e781fab6db6435876f5597https://doi.org/10.48550/arxiv.2402.14760
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs2024
  2. 2Fine-Tuning Language Models with Reward Learning on Policy2024 · 1 citations
  3. 3Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model2024
  4. 4Aligning Crowd Feedback via Distributional Preference Reward Modeling2024 · 1 citations
  5. 5Distributionally Robust Reinforcement Learning with Human Feedback2025