Randomized trial demonstrates enhanced verification of agent-written applications, suggesting greater reliability in software delivery.
AI coding agents can generate implementation changes faster than humans can verify them, making trustworthy verification the limiting resource in software delivery. Expectation-driven development (EDD) addresses that imbalance with an independently governed machine oracle whose central mechanism is Gold-State Parity™: a trusted, fixed, versioned gold state supplies both the scenario inputs and the expected persisted result. The test store is never seeded with gold: it begins from a declared baseline, receives inputs only through a declared application entry path, and must reproduce gold's expected-state projection under an explicit, fail-closed parity contract. One corpus supplies stimulus and expectation without circularity. The standing gate has two lanes. Proof drives a coverage-selected subset through the real user interface, mechanically evaluates the declared interface obligations, establishes browser-to-persistence state parity, and captures the normalized command-contract variants the UI emits. Validate instantiates only those captured variants with the full in-scope corpus through the application's official command entry point, tethered to the current Proof evidence by identity and an out-of-band provenance witness. A passing gate establishes contract-exact observational equivalence at the declared persistence boundary for the identified build, runtime profile, and oracle version, and nothing outside the versioned coverage model, which is the principal residual limitation and matures across immutable oracle versions. A failing gate is engineered to be as consumable as a passing one: a covered disagreement surfaces as a mechanical, localized artifact naming tables, business keys, and field-level mismatches, built to feed a diagnose-repair-rerun loop, whether agent-operated or human-operated, rather than to be reconstructed from a diff. A repair loop can only heal to the standard of the verdict that stops it: a loop stopped by interface tests heals until the screens behave; a loop stopped by Gold-State Parity™ heals only until what was entered is what is stored. The companion field report exercises that difference — seven seeded create-path mapping faults passed every interface-level test, and no interface, contract, validation, or read-path check convicted any of them, while the full-row comparison against gold supplied all seven gate convictions. The faults were all of one deliberately chosen kind, and the author reports his own run: the result shows what only the comparison could catch, not how often the gate catches faults in general. Read-path rendering, security, performance, concurrency, and external delivery remain separate gates. The component techniques have substantial prior art; the contribution claimed here is their composition, constraints, and governance as a standing acceptance gate for agent-written stateful write paths and, through the gate's engineered failure artifact, as the repair infrastructure that makes delegating the fix itself a governed act. The construction is exercised in a reference implementation that conforms to the specification as published (§9). In plain terms. The method exists to make done demonstrable. It fits an application whose business state lives in a store that can be copied, frozen, and compared, written through one or more official entry points: one corpus then supplies both the question and the answer — the values a scenario enters and the exact state the application must have stored — so done stops being a claim someone makes and becomes a thing the system shows. Because a refusal names the record and the field that disagreed, it is precise enough to repair against, by a human or by an agent, under a verdict the implementing agent cannot edit. The gate buys that precision by claiming only what it covers: a gate that claimed everything would be worth nothing.
No takes yet. Share an insight, caveat, or question.
Melvin Fahnestock (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: