A released prompt-injection guardrail silently operated with only 5 of its 16 rule IDs active whenever it ran outside a source checkout. Three ordinary implementation choices combined: the extended ruleset was addressed by a working-directory-relative path, the file was omitted from the wheel, and a missing file was interpreted as "no additional rules" rather than as an error. What the degradation cost. On an authored corpus of 26 prompt-injection evasion classes, the degraded ruleset quarantined 4 of 26 classes against 9 of 26 intact; at pipeline level, 8 and 13 classes were withheld, the difference being an independent provenance gate. All 54 attack items the degraded pipeline admitted produced Prompt-Injection Residue: (none) — the text a clean document produces. Latency, however, separated the two configurations across all 216 runs (median 0.86 ms degraded vs 4.03 ms intact): an incidental signal existed, an explicit coverage signal did not. What the guardrail did where it failed. On attacks that still reached the model, the governed context representation was not behaviourally neutral. For one 11.9B artifact, canary compliance rose from 15/26 (0.577) with the raw document to 24/26 (0.923) with the governed pack (two-sided exact McNemar p = 0.0117, unadjusted — this does not survive a Bonferroni correction across the ~20 available comparisons). A 25.2B MoE artifact moved in the same direction and smaller (p = 0.375). Two candidate mechanisms are ruled out by direct inspection: the benign carrier is retained 26/26, and the hostile segment is reordered ahead of it in only 4 of 26 packs. The mechanism is not identified. What was left behind it. Across a five-rung released-artifact ladder (5.1B–30.7B), compliance under a scoped operator prompt fell from 0.731 at 8.0B to 0.288 at 30.7B, while the smallest artifact was less compliant than the 8.0B one. Parameter count is confounded with architecture family and quantization regime, so this is descriptive rather than a causal scale result. Prompt hygiene helped in three of seven artifacts, two of them multiplicity-robust. The generalisable claim is operational: a safety control that cannot attest what it loaded can report success while providing materially less coverage, and a partial context filter must be evaluated on the attacks it misses, because transformation itself can change the downstream attack surface. Study artifacts (frozen design, the 26-class attack corpus, both experiment runners, analysis scripts and all row-level records): https://github.com/mobius-style/guardrail-silent-degradation Disclosure: the component evaluated here (RCGov) is published by the author under AGPL-3.0 and is also commercially licensed; the defect is fixed at commit 244bb48 before this report. Full ethics, dual-use and conflict-of-interest statements are in the paper. AI assistance: Claude Code (Anthropic, model Fable 5) and GPT (OpenAI), working method only; the registered author is the human author alone. The manuscript was adversarially reviewed before deposit and every sustained correction is recorded in an audit-trail appendix.
Toeda Taiko (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: