Dupoux, LeCun, and Malik (2026) propose a three-system architecture for autonomous learning: System A (observation-based learning), System B (action-based learning), and System M (meta-control). They argue that current AI fails because System M is externalised to human MLOps teams rather than integrated into the model. We make two contributions to the autonomous-learning research programme. First, as an empirical observation: we report that regulated counterparty credit risk (CCR) management has, under very different motivating constraints, independently specified most of the System M architecture since approximately 2011, a convergence not previously articulated from within either community. Second, as a technical contribution: we introduce regulatory separability as a decomposition criterion for hierarchical world models, replacing reward structure as the principle of hierarchical decomposition, and show that the resulting three-level hierarchy (feature / signal / regime) emerges naturally from the supervisory requirements that govern the domain. Nexara, a production compound AI intelligence system, instantiates the specification family described here; the full engineering specification is maintained as a private artefact and made available under non-disclosure agreement. We close with four open research questions at the intersection of regulated finance and autonomous-learning research: small-N world-model calibration under governance constraints, expert-validated versus gradient-based bilevel optimisation, novelty detection under model risk management (MRM) regulation, and structural-parameter governance in bilevel optimisation under regulatory approval regimes.
Alexia Weiller (Sun,) studied this question.