Comparison of AI coding tool productivity reveals gaps in vendor claims and actual outcomes, suggesting caution for organizations.
AI coding tools and "harness engineering"—the practice of designing constraintinfrastructure around AI agents—represent the latest iteration of a managementfashion cycle (Abrahamson 1996) that has recurred from CASE tools through Agile,DevOps, and microservices. Applying the four-dimensional matching frameworkdeveloped in a companion paper (Sophia 2026a), this paper presents a systematiccomparison of vendor-sponsored and independent empirical evidence on AI codingtool productivity. Vendor-affiliated studies consistently find 21–56% improvements onisolated, greenfield tasks using activity metrics; independent studies measuringsystem-level outcomes in mature codebases find neutral-to-negative results, with arandomized controlled trial showing experienced developers 19% slower with AI toolsdespite believing themselves 24% faster—a 43-percentage-point perception-realitygap that calls into question all self-report evidence. The paper documents theenterprise reality gap: the prerequisites for effective AI tool use—reliable automatedtesting, well-structured architecture, mature CI/CD, and sufficient engineeringcapability—are met by an estimated 1–3% of organizations. In the remaining 97–99%,AI tools function not as productivity enhancers but as vulnerability amplifiers,accelerating structural decay, propagating security vulnerabilities, degrading reviewcapacity, and eroding developing capability. The current cycle introduces a structuralnovelty unprecedented in prior fashion cycles: the entities promoting the constrainingdiscipline are the same entities selling the tools whose unreliability necessitates theconstraints.
No takes yet. Share an insight, caveat, or question.
Franny Philos Sophia (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: