PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 7, 20251 citationsOpen Access

Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning

View Full Paper
WYWon Sik YangSMShuming MaYLYankai Lin

Key Points

  • Longer chains of thought can impair reasoning performance of large language models in certain tasks.
  • An optimal scaled length distribution for reasoning efforts exists and differs across domains.
  • Using a small seed data set, the model learns to adapt reasoning efforts for better performance.
  • Self-improved models built on specific architectures outperform others across various mathematical benchmarks.

Abstract

Recent studies have shown that making a model spend more time thinking through longer Chain of Thoughts (CoTs) enables it to gain significant improvements in complex reasoning tasks. While current researches continue to explore the benefits of increasing test-time compute by extending the CoT lengths of Large Language Models (LLMs), we are concerned about a potential issue hidden behind the current pursuit of test-time scaling: Would excessively scaling the CoT length actually bring adverse effects to a model's reasoning performance? Our explorations on mathematical reasoning tasks reveal an unexpected finding that scaling with longer CoTs can indeed impair the reasoning performance of LLMs in certain domains. Moreover, we discover that there exists an optimal scaled length distribution that differs across different domains. Based on these insights, we propose a Thinking-Optimal Scaling strategy. Our method first uses a small set of seed data with varying response length distributions to teach the model to adopt different reasoning efforts for deep thinking. Then, the model selects its shortest correct response under different reasoning efforts on additional problems for self-improvement. Our self-improved models built upon Qwen2.5-32B-Instruct outperform other distillation-based 32B o1-like models across various math benchmarks, and achieve performance on par with QwQ-32B-Preview.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2025) studied this question.

synapsesocial.com/papers/68e585d0b1e78cc4e5f464f6https://doi.org/10.48550/arxiv.2502.18080
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling2025
  2. 2Thought calibration: Efficient and confident test-time scaling2025 · 1 citations
  3. 3Dynamic Early Exit in Reasoning Models2025 · 1 citations
  4. 4Concise thoughts: Impact of output length on LLM reasoning and cost2026 · 4 citations
  5. 5Fractured Chain-of-Thought Reasoning2025