Generative artificial intelligence has rapidly moved from experimental code completion to an increasingly integrated component of software development workflows. However, evidence concerning its effect on engineering productivity remains heterogeneous. Controlled experiments have reported substantial acceleration on bounded implementation tasks, whereas recent field experiments involving experienced developers working in familiar repositories have observed productivity slowdowns under AI-assisted conditions. At the same time, developer surveys and security studies identify persistent concerns regarding output accuracy, debugging effort, insecure code, and the verification of machine-generated changes. This paper presents a structured evidence synthesis of empirical research on generative AI in software engineering. We examine four questions: how AI assistance affects developer productivity; why reported productivity effects differ across experimental settings; what evidence exists regarding verification, security, and repository-scale limitations; and what engineering controls follow from the combined evidence. The synthesis indicates that the productivity effect of AI coding assistance is strongly context dependent. Benefits are most consistently observed for bounded and well-specified generation tasks, while evidence from mature repositories highlights the importance of implicit requirements, repository familiarity, and human validation. Security studies further demonstrate that functional plausibility does not guarantee secure implementation. Based on the synthesized evidence, we propose Continuous Verification and Sandboxing (CVS), a reference architecture that places isolated execution, static analysis, security scanning, and test-based evaluation between machine generation and human peer review. CVS is presented as an evidence-derived design implication rather than an empirically validated intervention. We conclude that generative AI should be evaluated as part of a socio-technical software delivery pipeline in which generation cost, verification cost, and integration risk are measured jointly.
Md.Nimur Rahman Durjoy (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: