PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 26, 20254 citationsOpen Access

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

View Full Paper
YYYuan YuanTST. SriskandarajahABAnna-Luisa Brakman

Key Points

  • Safe-completion training enhances model helpfulness and safety for language models, reducing the risk on dual-use prompts.
  • The approach helps mitigate binary classification issues, especially in complex scenarios with obscured user intent.
  • Incorporated into GPT-5, the training demonstrated significant improvement in safety outcomes through controlled experiments.
  • The findings suggest that prioritizing output centrality in safety can lead to better performance and reduced severity of failures.

Abstract

Large Language Models used in ChatGPT have traditionally been trained to learn a refusal boundary: depending on the user’s intent, the model is taught to either fully comply or out-right refuse. While this is a strong mitigation for explicitly malicious prompts, focusing safety training on refusals can lead to brittleness for prompts with obscured user intent. Binary refusal boundaries are especially ill-suited for dual-use cases (such as biology or cybersecurity), where a user request can be answered safely at a high level, but in some cases can lead to malicious uplift if sufficiently detailed or actionable. As an alternative, we propose safe-completions: a safety-training approach that centers on the safety of the assistant’s output, rather than a bi-nary classification of the user’s intent. Safe-completions seek to maximize helpfulness within the safety policy’s constraints. We incorporated this approach into GPT-5 and find that across both production comparisons and internally controlled experiments, safe-completion training improves safety (especially on dual-use prompts), reduces the severity of residual safety failures, and substantially increases model helpfulness.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yuan et al. (2025) studied this question.

synapsesocial.com/papers/68d6cd68b1249cec298b3a5bhttps://doi.org/10.70777/si.v2i6.15625
Ask AI
Helpful
Bookmark
Share
View Full Paper