This theoretical framework explores decision-making efficiency in AI systems, suggesting minimal internal transparency for optimal agency.
The Opacity Hypothesis proposes that optimal artificial agency does not require maximal internal transparency. This paper introduces the concept of Structural Opacity: global metacognition is not computationally free, and integrating high-dimensional internal states (ℝᴺ) without conditional independence assumptions scales super-linearly, approaching O(N²) overhead in worst-case dense dependency graphs. In resource-bounded systems, this creates strong selective pressure toward dimensionality reduction, projecting micro-dynamics into a low-dimensional (ℝᵏ) heuristic interface. We argue that operating through this simplified macro-state enables robust decision-making in non-stationary, out-of-distribution (OOD) environments. We propose a provisional functional metric (Ω) capturing the trade-off between information loss and OOD performance, and formalize how this compressed interface can act as a structural buffer against catastrophic forgetting in Continual Reinforcement Learning. Finally, we explicitly decouple this architectural framework from the metaphysics of phenomenal consciousness, focusing strictly on the algorithmic constraints of autonomous agency, and outline testable behavioral implications for future work on instrumental convergence and AI safety.
No takes yet. Share an insight, caveat, or question.
Lorenzo Gennari (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: