Experimental evaluation reveals limited behavioral efficacy of learned corrections in language models, indicating that formal authority guarantees diverge from semantic success.
How can a limited human correction remain consequential within a system whose computation exceeds the information contained in that correction? I develop this question through the receiver-relative framework of Self-Aware Networks and distinguish five obligations: legitimate authority, preserved meaning, effective implementation, measured consequences, and justified recovery. A versioned reference application tests scoped commitments through delegation, revocation, provenance-sensitive actions and learning. Its formal guarantees concern an explicitly mediated action boundary, not the competence or values of the proposing model. Experiments progress from controlled learned networks and GPT-2 to a pinned quantized Qwen3.5-0.8B hybrid decoder. In the principal learned-correction study, all 48 ordinary held-out answers are correct. Linear and nonlinear attention-state predictors each achieve only 5 of 32 requested reversed-role answers and no complete bidirectional story group, despite preserving all 16 tested color answers. Privileged corrected-state controls succeed in all eight groups. An authorized learned edit can therefore remain behaviorally defective, and stopping further edits is distinct from restoring the preceding state. These findings motivate separate acceptance criteria for authority, semantic success and recovery. They do not demonstrate scalable human oversight, autonomous understanding of values, frontier-model generalization or consciousness. The contribution is an executable and mathematically bounded account of how guidance can remain effective, and where that account still fails. Preprint v1.0; integrated Draft 10, September 5, 2026. The attached shared companion includes the applications, protocols, data, figures, 31 accepted Lean statements under recorded assumptions, and a later full fixed learned-Qwen pipeline reproduction (49 serial steps; 465 original answers and 465 separate-code reconstructions). The same experiments support both AI papers; these are not independent replications. The negative learned-correction result is retained. Authoring-agent checks and same-workflow reconstruction are not independent scientific review. Base-model and runtime binaries and private source originals are not included. License scope: CC BY 4.0 for original manuscript prose, figures, documentation and data. Original code is public for review without a separate reuse license; third-party materials retain their own terms. See the companion LICENSE.md and notices.
No takes yet. Share an insight, caveat, or question.
Micah Blumberg (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: