Large language models (LLMs) have recently shown promising potential in automating unit test evolution for evolving software systems. However, the effectiveness of LLMs in unit test evolution remains insufficiently understood, particularly with respect to prompt design choices, in-context learning (ICL) strategies, and different types of test evolution. In this paper, we present the first comprehensive empirical study to evaluate LLMs for unit test evolution. We systematically assess nine open-source code LLMs (3B to 34B parameters) and three state-of-the-art commercial models across diverse prompt designs, ICL strategies, and representative test evolution frameworks. To support robust and execution-based evaluation, we construct a new benchmark consisting of 530 real-world focal method–test co-evolution instances collected from seven actively maintained open-source projects. Our evaluation employs a suite of compilation, execution, and coverage-based metrics. Extensive experimental results reveal that prompt design and ICL methods significantly impact LLM effectiveness. Furthermore, the optimal configurations of these strategies vary substantially across different LLMs and evolution types. Based on our findings, we derive actionable insights to guide future research and practical adoption of LLM-based techniques for unit test evolution.
No takes yet. Share an insight, caveat, or question.
None et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: