Randomized trial evaluates compliance outcomes in conversational AI, indicating improved governance effectiveness.
Large Language Models are increasingly deployed in domains where conversations are subject to legal, regulatory, or organizational constraints. Existing approaches primarily address these challenges through model alignment, safety training, or response filtering. While these techniques have significantly improved the safety of language generation, governance itself remains largely embedded within the behavior of the underlying model, limiting transparency, auditability, and independent evaluation. This report introduces Interactive Runtime Governance (IRG), a computational architecture that treats governance as an explicit runtime capability rather than an implicit property of language generation. The proposed framework separates semantic evidence extraction, deterministic governance adjudication, explanatory user guidance, compliance verification, and behavioral evaluation into independent runtime components. By decoupling governance from the language model, IRG enables governance policies to be specified, inspected, modified, and empirically evaluated independently of the underlying foundation model. Beyond the architectural contribution, this report proposes a behavioral evaluation methodology that measures governance through observable conversational outcomes rather than response quality alone. Instead of asking whether a language model generated an acceptable response, the framework evaluates whether governance successfully guided conversations toward compliant and productive interaction. To support future empirical validation, the report further defines a pre-registered experimental protocol, identifies potential threats to validity, and discusses broader implications for runtime governance in regulated conversational AI. The central hypothesis of this work is that governance should be understood as an independent computational discipline operating alongside, rather than within, language generation. While empirical validation remains future work, the proposed architecture establishes a formal foundation for transparent, explainable, and measurable runtime governance that complements existing advances in alignment and AI safety.
No takes yet. Share an insight, caveat, or question.
Teodor Minchev (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: