Benchmark evaluates task recovery in LLM agents, suggesting a verification-first design rule.
LLM agents in long-horizon workflows often fail not through visible crashes but through accumulated task-state inconsistencies, including memory conflicts, goal drift, false premises, contradictions, and repair-induced cascades. Standard evaluations measure final task success, but production reliability also requires a narrower capability: when task state becomes inconsistent, can an agent detect the failure, repair it, and verify recovery? We present CRepair, a benchmark and runtime intervention for this detect-repair-verify problem, consolidating four pilot studies. CRepair scores repair-loop closure as C_repair = D x R x V x S across 13 scenarios and six structural failure types. Across a benchmark study, a runtime-wrapper pilot, a three-model replication using Claude Sonnet, Gemini 2.5 Flash, and GPT-4o, and an eight-condition ablation, a consistent pattern emerges: explicit verification is the load-bearing component of structured self-repair. A lightweight wrapper enforcing targeted repair and verification improves C_repair in all 9 runs, while generic retry is indistinguishable from baseline. Verification-containing scaffolds occupy the top performance tier. Verify-only yields Delta C = +0.208, with paired t = 4.01 across runs. Detection-only scaffolding reduces coherence below baseline, with Delta C = -0.236 and t = -3.44, and is negative in all runs, consistent with prompt-induced over-identification. Whether verification alone matches or exceeds the full four-step loop cannot be resolved at this sample size; the two are statistically indistinguishable. All results are pilot-scale and depend on LLM-as-judge scoring pending human validation. The practical conclusion is a verification-first design rule: require explicit repair-loop closure before accepting a recovered output, and never deploy detection prompts without a downstream repair or verification path.
No takes yet. Share an insight, caveat, or question.
Kaminovs Sergejs (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: