Randomized trial demonstrates cheaper, safer, and auditable multi-agent work in AI systems under human authority.
Large-language-model (LLM) agents are increasingly used for multi-step knowledge work, but the prevailing way of running and coordinating several of them — as long-lived processes that poll for work and hold their state in an ever-growing context window — is expensive, forgetful, and unsafe at the boundaries that matter. We present an architecture for event-driven coordination of episodic AI agents that decouples an agent's durable identity and memory from any running process. Each agent exists as a sequence of bounded episodes separated by dormant intervals that consume no inference compute; an episode begins only when a durably-recorded event addressed to the agent arrives, at which point a compute embodiment is instantiated and the agent's prior working state is rehydrated from a compact, self-authored memory digest. Four mechanisms distinguish the design: (1) state changes and the events that describe them are committed atomically in one transaction, so events are never lost relative to, nor observable before, the state they describe; (2) a near-zero-cost wait/notify mechanism with per-agent cursors wakes an agent only when an addressed event arrives, giving idempotent, exactly-once handling; (3) typed request and task state machines govern communication and work, including a human-only verification gate enforcing the invariant that an agent never certifies its own work; and (4) a graduated, reversibility-keyed authority model executes reversible actions autonomously while hard-gating irreversible or high-consequence actions to a human, with every decision audited. The coordination core is domain-agnostic. We describe a working reference implementation, Orcha (open-source, MIT, openorcha.io), and report a reflexive case study in which the system coordinated agents building and releasing its own software under human authority. We argue the architecture makes multi-agent work cheaper while idle, coherent across unbounded sessions, safe at irreversible boundaries, and auditable end-to-end.
No takes yet. Share an insight, caveat, or question.
Kedar Ravindra Haldankar (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: