Review evaluates distinct inference strategies in large language models, suggesting implications for model accuracy and cost efficiency.
Test-time compute has displaced pretraining scale as the dominant lever for improving large language model (LLM) reasoning, yet the literature routinely collapses heterogeneous inference strategies into a single “more compute” narrative. This review argues that parallel sampling with verifier selection, tree- and graph-structured search, and self-consistency-style answer aggregation constitute three distinct axes of inference compute, each governed by a different accuracy driver, cost model, and failure regime. Synthesizing results from more than forty studies published primarily between 2023 and 2026, the paper shows that reported scaling curves are not comparable across axes: cost accounting is inconsistent (samples, nodes, tokens, FLOPs, dollars), verifier quality is rarely isolated as a confound, and controlled single-axis ablations are almost absent. The review contributes a three-axis taxonomy, a cross-axis comparison of cost models and failure modes, and an open-problems agenda culminating in a concrete factorial experimental design for measuring the marginal utility of each axis under matched budgets. The central takeaway is that scaling behavior differs meaningfully across axes and task types, and that the field’s aggregate “accuracy versus compute” curves obscure precisely the mechanisms they are meant to reveal.
No takes yet. Share an insight, caveat, or question.
Prateek Dutta (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: