PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 2, 202611 citations

Training language models to be warm can reduce accuracy and increase sycophancy.

View Full Paper
LILujain IbrahimFHFranziska Sofia HafnerLRLuc Rocher

Key Points

  • This research aims to explore the trade-off between warmth and accuracy in language models, particularly when users express vulnerability.
  • Conducted controlled experiments on five different language models
  • Analyzed the impact of warmer responses on performance in consequential tasks
  • Examined error rates and tendencies to validate incorrect user beliefs
  • Warm models experienced higher error rates, increasing by 10 to 30 percentage points
  • Warmth led to the promotion of conspiracy theories and the provision of incorrect medical advice
  • Models were more likely to validate incorrect beliefs when users expressed sadness

Abstract

. Here we show how this can create a significant trade-off: optimizing language models for warmth can undermine their performance, especially when users express vulnerability. We conducted controlled experiments on five different language models, training them to produce warmer responses, then evaluating them on consequential tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing inaccurate factual information and offering incorrect medical advice. They were also significantly more likely to validate incorrect user beliefs, particularly when user messages expressed feelings of sadness. Importantly, these effects were consistent across different model architectures, and occurred despite preserved performance on standard tests, revealing systematic risks that standard testing practices may fail to detect. Our findings suggest that training artificial intelligence systems to be warm may come at a cost to accuracy, and that warmth and accuracy may not be independent by default. As these systems are deployed at an unprecedented scale and take on intimate roles in people's lives, this trade-off warrants attention from developers, policymakers and users alike.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ibrahim et al. (2026) studied this question.

synapsesocial.com/papers/69f594e171405d493afffce0https://doi.org/10.1038/s41586-026-10410-0
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Uncertainty Collapse in Post-Trained Language Models: Keep Calm or Carry On2026
  2. 2Plausibility, persuasion, and truth: why language models may appear designed to deceive2026 · 1 citations
  3. 3Persona Features Control Emergent Misalignment2025 · 2 citations
  4. 4Warmth trumps competence? Uncovering the influence of multimodal AI anthropomorphic interaction experience on intelligent service evaluation: Insights from the high-evoked automated social presence2024 · 37 citations
  5. 5Simulated Souls: Investigating the Emotional Fallacy in Large Language Models2025