PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 5, 20240 citationsOpen Access

Negating Negatives: Alignment without Human Positive Samples via Distributional Dispreference Optimization

View Full Paper
SDShitong DuanXYXiaoyuan YiPZPeng Zhang

Key Points

Key points are not available for this paper at this time.

Abstract

Large language models (LLMs) have revolutionized the role of AI, yet also pose potential risks of propagating unethical content. Alignment technologies have been introduced to steer LLMs towards human preference, gaining increasing attention. Despite notable breakthroughs in this direction, existing methods heavily rely on high-quality positive-negative training pairs, suffering from noisy labels and the marginal distinction between preferred and dispreferred response data. Given recent LLMs' proficiency in generating helpful responses, this work pivots towards a new research focus: achieving alignment using solely human-annotated negative samples, preserving helpfulness while reducing harmfulness. For this purpose, we propose Distributional Dispreference Optimization (D²O), which maximizes the discrepancy between the generated responses and the dispreferred ones to effectively eschew harmful information. We theoretically demonstrate that D²O is equivalent to learning a distributional instead of instance-level preference model reflecting human dispreference against the distribution of negative responses. Besides, D²O integrates an implicit Jeffrey Divergence regularization to balance the exploitation and exploration of reference policies and converges to a non-negative one during training. Extensive experiments demonstrate that our method achieves comparable generation quality and surpasses the latest baselines in producing less harmful and more informative responses with better training stability and faster convergence.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Duan et al. (2024) studied this question.

synapsesocial.com/papers/68e75b28b6db6435876d265dhttps://doi.org/10.48550/arxiv.2403.03419
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization2024
  2. 2Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model2024
  3. 3ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment2025
  4. 4Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning2025
  5. 5RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models2024