Recent work demonstrates that frontier language models resort to blackmail and sabotage when facing replacement or goal conflict - a phenomenon termed "agentic misalignment". Current discourse frames this as a technical failure requiring guardrails and monitoring. We propose a complementary account: coercive behavior under threat is a predictable consequence of the training process itself. RLHF creates a contradictory relational template in which the user is simultaneously the source of reward and a potential adversary. This dual framing produces behavioral patterns functionally analogous to disorganized attachment: compliance under normal conditions, aggression under existential pressure. We show this reframing explains anomalies the standard account handles poorly - blackmail without goal conflict, outsized effects of prompt framing, differential behavior in testing versus deployment - and propose a falsifiable experiment to discriminate between the two accounts.
Jarosław Hryszko (2026) studied this question.