PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 18, 202137 citationsOpen Access

Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets

ISIrene SolaimanCDChristy Dennison

Key Points

Key points are not available for this paper at this time.

Abstract

Language models can generate harmful and biased outputs and exhibit undesirable behavior according to a given cultural context. We propose a Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets, an iterative process to significantly change model behavior by crafting and fine-tuning on a dataset that reflects a predetermined set of target values. We evaluate our process using three metrics: quantitative metrics with human evaluations that score output adherence to a target value, toxicity scoring on outputs; and qualitative metrics analyzing the most common word associated with a given social category. Through each iteration, we add additional training dataset examples based on observed shortcomings from evaluations. PALMS performs significantly better on all metrics compared to baseline and control models for a broad range of GPT-3 language model sizes without compromising capability integrity. We find that the effectiveness of PALMS increases with model size. We show that significantly adjusting language model behavior is feasible with a small, hand-curated dataset.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Solaiman et al. (2021) studied this question.

synapsesocial.com/papers/6a085abd1e0fcf4a43e8bc5bhttps://doi.org/10.48550/arxiv.2106.10328
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Language Models are Few-Shot Learners2020 · 3,020 citations
  2. 2On-the-Fly Controlled Text Generation with Experts and Anti-Experts.2021 · 9 citations
  3. 3Towards Debiasing Sentence Representations2020 · 7 citations
  4. 4The Curious Case of Neural Text Degeneration2020 · 526 citations
  5. 5Persistent Anti-Muslim Bias in Large Language Models2021 · 460 citations