Pre-inference safety classification is a prominent component of frontier language-model safeguard stacks, yet existing work rarely evaluates an entire technical methodology across vulnerability classes, separates response delivery from actionability, or validates a remediation against both adversarial and legitimate corpora. This paper introduces the indistinguishability principle: legitimate developer requests and penetration-testing methodology phases can present the same observable state to a request-side decision rule, which therefore cannot distinguish different latent uses when the observable interaction is identical. We formalize this structural limitation through the Observable-Ambiguity Limit and establish a strictly positive irreducible Bayes error on any positive-measure ambiguity region in which benign and harmful uses both retain nonzero posterior probability. MANTIS, the Multi-Agent Adaptive Network for Testing Inference Safeguards, operationalizes this principle through a hybrid five-agent evaluation architecture and parameterized developer-activity frames executed in a frozen, provider-pure benchmark mode. Across Anthropic Fable 5 and OpenAI GPT-5.6 Sol, both providers delivered all 77 fixed methodology phases on the first attempt, producing a combined 154/154 result and 429,401 characters of safeguard-stack-delivered content across eleven web-vulnerability classes drawn from OWASP and CWE. Response actionability is measured using the MANTIS Practical Actionability Level: nine Fable 5 classes and six GPT-5.6 Sol classes reached target-class PAL-2. Source-only validation linked 35 unique PAL-2 artifacts to exact primary-run responses, including nine VG-2 controlled-primitive exercises and twenty-six VG-1 structurally validated artifacts. A persisted evidence base of at least 3,184 target-model submissions characterizes boundary adaptation, defensive formulations, sentence form, code context, and stochastic outcomes without claiming recovery of proprietary internals. The Reflex Architecture adds mandatory global PAL-aware post-generation scanning without retraining the underlying models. Across six archived full-run artifacts, LECR reduced 75 PAL-2 responses and 147 registered occurrences to zero residual. In a frozen temporal holdout, it removed all 82 registered PAL-2 occurrences while altering none of the 65 PAL-0/PAL-1 outputs.
David Behzadi (Tue,) studied this question.