Benchmark evaluation reveals substantial discrepancies between reported success and functional preservation in LLM agents, indicating tool-dependent performance variations.
We introduce RefactorBench-JS, an open benchmark for evalu ating AI coding agents on behavior-preserving software maintenance tasks. The benchmark measures whether LLM agents can decompose large JavaScript and React files into smaller modules without chang ing observable behavior. It consists of 123 scored fixtures—spanning algorithmic logic, data modules, utility/API logic, web UI, and mobile UI—each paired with a hidden unit test suite: hidden from the agent during evaluation, then released publicly for auditability. An agent’s output is scored by executing the hidden tests against the refactored filesystem, directly measuring whether behavior was preserved. This design is grounded in a simple principle: behavioral preservation is a functional property, and functional properties are best verified by functional tests—not by code similarity metrics, static analysis, or LLM-as-judge. We demonstrate RefactorBench-JS by evaluating Anthropic Claude and Google Gemini models across two tool configurations. The benchmark reveals model-dependent tool effects, large gaps between agent-reported success and actual behavioral preservation, and dis tinct failure modes across model families. We report hidden-test pass rates, calibration, operational metadata, and a failure taxonomy. These findings informed model and tool selection in a production refactor ing system whose observed success rate improved from approximately 65% to above 97% amid multiple engineering changes; this production result is observational, not a controlled causal estimate. Refactor Bench-JS is publicly available at https://github.com/Create-Inc/ refactor-bench.
No takes yet. Share an insight, caveat, or question.
Daniel Tianming Chen (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: