Novel framework improves mathematical reasoning outcomes in intelligent systems, suggesting effective supervision methods.
Large language models (LLMs) are increasingly used in applied intelligent systems, but mid-sized models still lag on mathematical reasoning, partly because reliable step-level supervision is scarce. Many existing remedies rely on costly human annotation, stronger teacher models, or heavy training pipelines, which limits practical adoption. We propose VGPO-MCTS (Value-Guided Group-wise Policy Optimization over Monte Carlo Tree Search), a search-and-distillation framework that constructs reusable step-level supervision from datasets that provide only problems and final answers. VGPO-MCTS augments a frozen backbone with (i) a lightweight value model that scores candidate reasoning states formed by a reasoning prefix and its candidate next step, and (ii) a policy updated with parameter-efficient adaptation. During search, the value model guides tree expansion and selection, while verified outcomes are propagated backward to correct node utilities. The corrected search trees are then distilled into two complementary datasets: a value regression dataset for value learning and group-wise sibling candidate sets for GRPO-style policy optimization. Experiments on GSM8K and the MATH dataset with ChatGLM3-6B and SciGLM-6B show stable round-wise improvements in final-answer exact match under a lightweight adaptation setting. After three rounds of self-training, the proposed framework improves performance by about 6.3 percentage points on GSM8K and about 3.9 percentage points on MATH across the two backbones.
No takes yet. Share an insight, caveat, or question.
Wu et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: