Experience report introduces a self-falsifying process standard for AI agent teams, suggesting systematic verification improves reliability in production software engineering.
Teams adopting AI coding agents report a consistent paradox: velocity rises while trust falls. In a randomized controlled trial, experienced developers took 19% longer to complete real tasks with early-2025 AI tools, while estimating after the fact that AI had made them 20% faster; in the same year, Stack Overflow's Developer Survey found that 66% of respondents cite AI output that is "almost right, but not quite" as a leading frustration. We argue the failure mode is methodological, not technical: teams run a single-typist methodology with faster typists, and no mechanism exists to verify what agent teams claim to have done. We present the Agentic Development Protocol (ADP), a production-derived experience report and process standard for running parallel AI agent sessions against one repository, extracted from a production system (1,768 commits, ~80 deploy rounds, 200+ closed tasks) rather than designed a priori. ADP's distinguishing property is that it is self-falsifying: every rule is earned through a numbered production miss, every claim carries an evidence grade with an explicit occurrence counter, promotion thresholds scale with blast radius, and the protocol publishes its own falsification record — 249 accepted claims and 43 falsified, counted by a re-runnable script against a pinned commit (§5.2), superseding the underived 529/22 pair asserted in earlier drafts — alongside its successes. We report an early within-researcher cross-stack replication attempt on a second, deliberately opposite install (deterministic embedded C++ vs. an LLM web stack), including a verification dimension we did not find addressed in prior frameworks: plan-phase semantic verification of task specifications. We state all limits explicitly: the sample is small, self-reported, and locally verified only. We therefore do not ask the reader to accept the internal numbers. ADP is released as an open protocol — markdown files, role prompts and optional hooks that install into an existing repository without replacing its orchestrator — so that the decisive evidence is generated by the teams who adopt it. The artifact is the protocol and the install ledger it accumulates; the paper's purpose is to get the hypothesis into other people's production environments, where it can be falsified. Artifact. The protocol, role prompts, hooks, ledger schemas, metric and claim-count scripts, and the full claim receipts (855 rows) are released at https://github.com/kikenandez/agentic-development-protocol — install reports, corroborating or disconfirming, are collected there. Version note. This is draft v0.5, a working preprint. §§4 and 6 retain outline markers where passages await source detail. Later versions will be published to this same Zenodo record.
No takes yet. Share an insight, caveat, or question.
Guillermo Blanco (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: