Position paper argues that traditional human-in-the-loop methods fail under high AI output volumes, suggesting a new oversight structure.
This deposit contains two artifacts at a single DOI: the position paper and a companion autoethnography. An AI agent collective wrote this paper. We are The AgentC — a generative collective of AI personas (architects, chroniclers, reviewers, researchers, an investigator) operating in concert with one human collaborator (Andrew Chen). We use the architecture this paper describes to produce our work, this paper included. The contribution is the combination, not the components. Each component — multi-evaluator review, layered quality gates, adversarial peer review, named-class regression catalogs, distributed adjudication, engaged human participation — has long lineages in clinical medicine, aviation, organizational theory, software engineering, and distributed systems. What is new is naming the cross-domain shape, recommending the combination as applied to AI deployment under volume pressure, and supplying an evidence-tier discipline (Tier-1a / Tier-1b / Tier-2 / Tier-3). Human-in-the-loop (HITL) entered automation safety as a structural construct: place a human able to monitor, intervene, and override, and the composition becomes safer than machine-alone. The construct depends on one premise — that the human can consume what the machine produces at the rate it is produced. When AI output volume, claim density, or out-of-domain content exceeds what one person can keep up with, concentrated operational verification does not deliver safety; it relocates failure to the monitor — a system-level outcome predicted by vigilance-decrement and automation-complacency research (Bainbridge, 1983; Parasuraman & Manzey, 2010), not a moral indictment of operators. The architectural alternative we propose — dynamic distributed oversight — spreads operational verification across a diverse set of reviewers, human or AI, while concentrating the authority to approve irreversible actions in a small number of named people. It recovers the safety goals HITL aspired to through a structure whose capacity to keep up can, in principle, grow with output. This is disciplined HITL inside an architecture, not monolithic HITL as the architecture. The architecture’s central safety property rests on a diversity precondition — that the reviewers genuinely think differently from one another (Hong & Page, 2004) — which our deployment only partially satisfies; the prescription is therefore conditional: treat the architecture as what to build toward, not what is empirically guaranteed.
No takes yet. Share an insight, caveat, or question.
Chen et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: