PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 20260 citationsOpen Access

Improving reasoning at inference time via uncertainty minimisation

View Full Paper
NLNicolas LegrandKEKenneth EnevoldsenMKMárton Kardos

Key Points

  • This research aims to improve multi-step reasoning in large language models by minimizing uncertainty during inference.
  • Proposed a strategy that frames reasoning as uncertainty minimisation.
  • Utilized self-certainty as a metric for decision-making at each reasoning step.
  • Conducted experiments on datasets MATH500 and GSM8K with various model sizes.
  • Evaluated cross-linguistic performance to test the robustness of the method.
  • Thought-level self-certainty maximization outperformed greedy decoding techniques.
  • Achieved results matching or exceeding self-consistency with fewer tokens used.
  • Cross-linguistic tests showed strong transferability across languages.
  • Early decision-making in the reasoning process was linked to final accuracy improvements.

Abstract

Large language models (LLMs) now exhibit strong multi-step reasoning abilities, but existing inference-time scaling methods remain computationally expensive, often relying on extensive sampling or external evaluators. We propose a principled strategy that frames reasoning as uncertainty minimisation and operates at the level of individual thoughts rather than tokens. Our method selects, at each reasoning step, the continuation that maximizes the model's self-certainty, a metric computed from its internal predictive distribution. This approach achieves significant improvement with a small number of samples, relies exclusively on model-internal signals, and applies to open-ended questions as opposed to methods like majority voting. Experiments on MATH500 and GSM8K across multiple model sizes demonstrate that thought-level self-certainty maximization consistently outperforms greedy decoding and matches or exceeds self-consistency under comparable token budgets. Cross-linguistic evaluations further indicate that the method transfers robustly beyond high-resource languages. Furthermore, analysis of self-certainty dynamics reveals that correct reasoning trajectories converge early to stable paths, suggesting that early decisions, likely associated with the planning of the reasoning process, are predictive of final accuracy. Building on this result, we show that self-certainty maximisation applied to the early steps can explain most of the performance gain and provide a simple yet efficient inference-time scaling method.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Legrand et al. (2026) studied this question.

synapsesocial.com/papers/69b4fa6fb39f7826a300b245https://doi.org/10.48550/arxiv.2603.07159
Ask AI
Helpful
Bookmark
Share
View Full Paper