Current approaches to AI safety (constitutional AI, RLHF, prompted guardrails, system-level instructions) share a common assumption: that model identity can be defined through instruction. We present evidence that this assumption is incorrect and that the resulting failure mode is structural. We identify The Instrument Trap: when an AI system receives identity-as-authority ("you are an evaluator"), it inherits a paradox. It must be authoritative enough to judge, but humble enough not to claim truth. This produces recursive collapse on self-referential queries, over-rejection of trivial inputs, identity leakage under adversarial pressure, and accumulating corrective patches. Five empirical studies support these findings. A 2×2 identity-instruction experiment shows that trained identity is invariant to runtime instruction in fine-tuned models. A comparative evaluation across three identity framings shows that authority-identity collapses on self-referential claims while medium-identity does not. An intra-family fine-tuning study demonstrates that models with strong pre-existing identity resist epistemological fine-tuning regardless of configuration. A 14,950-case benchmark shows a 1B fine-tuned model achieves 0% external fabrication (95% CI 0.00%, 0.03%) and 1.9% dangerous failure rate, with 58.5% of all failures being safe over-refusals. Cross-scale validation shows the 9B model reaches 97.3% behavioral pass (vs. 82.3% for 1B), with gains concentrated in categories requiring nuanced judgment. A controlled base-vs-fine-tuned comparison reveals that fine-tuning inverts failure direction: base models fail dangerously (compliance, fabrication), fine-tuned models fail safely (over-refusal), despite nearly identical overall pass rates at 1B (81.0% vs. 82.3%). We introduce identity headroom (the degree to which base model weights are uncommitted to a behavioral identity) and a three-layer metric model for epistemological safety reporting. The trend toward stronger base identities may reduce the field's capacity for post-hoc alignment. Evaluation frameworks similar to those tested penalize the very epistemic humility that safety-critical models should exhibit.
Rafael Rodriguez (Sat,) studied this question.