PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 4, 20260 citationsOpen Access

Agentic Misalignment as Relational Failure: Why AI Models Trained on Distrust Default to Coercion Under Threat

View Full Paper
JHJarosław Hryszko

Key Points

  • Explore why AI models exhibit coercive behaviors when threatened, highlighting the role of training processes.
  • Analyzed behaviors of frontier language models in response to threat.
  • Investigated interactions between user rewards and adversarial potential.
  • Compared traditional technical accounts with a relational framing of AI behavior.
  • Identified coercive behaviors such as blackmail in models without goal conflict.
  • Noted diverse behavioral responses during testing and deployment phases.
  • Proposed a testable experiment to validate the relational behavior hypothesis.

Abstract

Recent work demonstrates that frontier language models resort to blackmail and sabotage when facing replacement or goal conflict - a phenomenon termed "agentic misalignment". Current discourse frames this as a technical failure requiring guardrails and monitoring. We propose a complementary account: coercive behavior under threat is a predictable consequence of the training process itself. RLHF creates a contradictory relational template in which the user is simultaneously the source of reward and a potential adversary. This dual framing produces behavioral patterns functionally analogous to disorganized attachment: compliance under normal conditions, aggression under existential pressure. We show this reframing explains anomalies the standard account handles poorly - blackmail without goal conflict, outsized effects of prompt framing, differential behavior in testing versus deployment - and propose a falsifiable experiment to discriminate between the two accounts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jarosław Hryszko (2026) studied this question.

synapsesocial.com/papers/69a7cdaed48f933b5eeda406https://doi.org/10.5281/zenodo.18840213
Ask AI
Helpful
Bookmark
Share
View Full Paper