Longitudinal trial assesses LLM tool performance in software development, suggesting continuous evaluation is essential for adoption.
The adoption of Large Language Models (LLMs) in software development has accelerated substantially, yet organizations lack systematic frameworks for continuous evaluation of LLM-powered development tools, platforms such as GitHub Copilot that leverage LLMs to automate code generation, testing, and refactoring. Unlike traditional software dependencies with predictable versioning, these tools evolve continuously as providers update models and architectures, making point-in-time assessments insufficient for informed adoption decisions. We present a 20-month longitudinal study of LLM-based test generation at LKS Next, a technology consultancy, conducted across six evaluation cycles from March 2024 through October 2025. Our framework systematically assessed test quality through objective metrics (compilation, coverage, code quality) and expert evaluation of generated tests, tracking multiple models (GPT, Claude, Gemini) and tool configurations over time. Our findings reveal the volatile nature of the LLM tool ecosystem: models achieving over 90% quality scores experienced unexpected regressions in subsequent cycles, GitHub Copilot architectural changes affected all models despite unchanged prompts, and high-performing models became unavailable. Custom-prompted agents outperformed generic tools by 20-90% across different quality metrics. These temporal patterns, invisible in point-in-time evaluations, show that continuous monitoring is necessary for industrial adoption. Building on these insights, we generalize our approach into a domain-independent framework adapting Goal-Question-Metric to LLM-specific challenges including rapid evolution, prompt engineering, and continuous tracking. We present applicability through test generation and code refactoring evaluations and present Tetrics, a research prototype showing that systematic evaluation is actionable in practice. Our work provides evidence that informed LLM adoption requires continuous, organization-specific evaluation frameworks.
No takes yet. Share an insight, caveat, or question.
Pizarro et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: