Systematic review reveals unaddressed validity threats across autonomous LLM agent benchmarks, highlighting an urgent need for standardized evaluation infrastructure.
Large language models (LLMs) are increasingly deployed as agents that perceive an environment, plan, invoke tools, and act over multi-step trajectories, and a fast-growing literature proposes benchmarks to evaluate them. We argue that the resulting scores are not self-interpreting: what a benchmark number means depends on which capability it measures, how that capability is scored, and where the agent is tested. This article moves beyond the prevailing benchmark catalogue to a validity-centered analysis of agent evaluation, built on a reproducible systematic mapping review under the PRISMA-ScR standard (259 primary studies; 294 in the released living-review corpus). The corpus is a carefully constructed, reproducible mapping sample—not a census of the field: primary-arm records were screened by a single audited automated pass (human double-screening, \(κ =0.65\) , was applied to the supplementary arm), and broad-map attribute shares are provisional heuristic estimates. The article makes four contributions. An operational definition—the dependent-step test—demarcates autonomous agentic evaluation from static natural-language-processing evaluation and prompt engineering. A meta-taxonomy locates every benchmark in a three-pillar coordinate system of capability ( what ), scoring paradigm ( how ), and environment topology ( where ). An evidence map over the corpus is paired with a purposive , hand-verified, venue-verified landscape matrix of seventeen prominent benchmarks (inter-rater reliability \(κ =0.77\) ). Finally, a critical analysis shows that data contamination, non-determinism, and execution cost form a structural trilemma: among these seventeen prominent, verified benchmarks, none reports evidence that the three threats are jointly controlled, and none reports a complete standardized run-cost record (0/17). A five-direction roadmap reframes progress as a shift from individual benchmarks toward standardized, reliability-first, cost-aware evaluation infrastructure, supported by an open companion repository.
No takes yet. Share an insight, caveat, or question.
Nageshwaran et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: