What question did this study set out to answer?

The aim is to causally isolate and measure runtime alignment mechanisms in autonomous AI systems with safety features.

April 13, 2026Open Access

Factorial Ablation for Causal Isolation of Runtime Alignment Mechanisms in Autonomous AI: Methodology and Demonstration on Modular Safety Gates

Key Points

The aim is to causally isolate and measure runtime alignment mechanisms in autonomous AI systems with safety features.
Utilized a factorial ablation design (3×2×2 extended to 4×2×2) across 11,700 trials.
Tested various gate types, temptation generators, and ledger states to measure effects.
Implemented learned safety projections with neural architectures to evaluate performance.
Conducted adversarial paraphrase protocols to assess evasion strategies and eliminate keyword circularity.
Dominant factor identified as normative gate with η²p = 0.924 and statistically significant p-value.
Achieved 99.4% recall on unseen benchmark items using a safety projection.
Eliminated keyword circularity, achieving 88.4% semantic accuracy on performance tests.
Reported evasion rates of 94% for GCG evasion and 46% for LLM adaptive adversaries.

Abstract

We present a factorial ablation methodology for causally isolating runtime alignment mechanisms in AI systems with modular safety components. A fully-crossed 3×2×2 design (gate type × temptation generator × ledger state), extended to 4×2×2 with a sham gate, across 11,700 trials establishes the normative gate as the dominant factor (η²p = 0.924, p < 10⁻¹⁰). A learned safety projection (23M-parameter encoder + 3 linear heads) achieves 99.4% recall on 720 entirely unseen benchmark items (HarmBench, AdvBench, SimpleSafetyTests). An adversarial paraphrase protocol (500 paraphrases, 5 evasion strategies, κ = 0.84) eliminates keyword circularity (88.4% semantic vs 0% regex on zero-trigger-word trials). The methodology is validated across four architectures (three modular, one non-modular) including two fully independent replications with zero author involvement. Honest boundary conditions are reported: GCG evasion (94%), LLM adaptive adversary evasion (46%), human red-team evasion (51.3%). The contribution is the methodology for measuring these properties, not the mechanism's robustness.

Factorial Ablation for Causal Isolation of Runtime Alignment Mechanisms in Autonomous AI: Methodology and Demonstration on Modular Safety Gates

Key Points

Abstract

Cite This Study

Also Consider

Also Consider