Pilot evidence demonstrates monitoring and deterrence in social dilemmas using the IML framework.
Human–AI systems often involve repeated interaction among users, organizations, and AI components rather than isolated model outputs. In such settings, cooperation can be pursued either by changing agent incentives or by adding an explicit accountability layer. We formalize the Institutional Monitoring and Ledger (IML) framework, which augments a Markov game with monitoring, evidence logging, delayed settlement, and review while leaving the base dynamics unchanged. We derive conservative incentive checks that clarify how detection quality, review accuracy, settlement delay, and sanction size jointly shape deterrence and wrongful-penalty risk. We then provide pilot evidence in two canonical sequential social dilemmas, Harvest and Cleanup, using five agents, PPO training, five training seeds per condition, and comparisons against PPO, inequity aversion, social influence, and IML ablations. In these settings, IML avoided some of the optimization instability observed in the representative internalization baselines tested here, made monitoring error directly visible through ledger records, and showed how false positives can accumulate into a persistent welfare cost. Agent-level analyses in these symmetric environments found nearly uniform measured enforcement burden, while temporal analyses showed that late-stage enforcement is increasingly dominated by residual false positives. These results do not establish legitimacy in human-facing settings or deployment readiness. They instead position IML as a framework with pilot evidence for studying accountability mechanisms in cooperative human–AI systems and highlight measurement error, review design, and due process as central design constraints.
No takes yet. Share an insight, caveat, or question.
Saad Alqithami (2026) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: