PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 8, 20250 citationsOpen Access

Think When You Need: Self-Adaptive Chain-of-Thought Learning

View Full Paper
JYJunjie YangKLKe LinYXYu Xing

Key Points

  • Adaptive learning enhances efficiency in language models with improved conciseness.
  • Evaluation across reasoning benchmarks demonstrates effectiveness in reducing response length.
  • The approach utilizes rewards based on performance rather than penalizing longer reasoning.
  • Supports more accurate problem-solving in scenarios lacking clear ground truth, highlighting its significance.

Abstract

Chain of Thought (CoT) reasoning enhances language models' performance but often leads to inefficient "overthinking" on simple problems. We identify that existing approaches directly penalizing reasoning length fail to account for varying problem complexity. Our approach constructs rewards through length and quality comparisons, guided by theoretical assumptions that jointly enhance solution correctness with conciseness. Moreover, we further demonstrate our method to fuzzy tasks where ground truth is unavailable. Experiments across multiple reasoning benchmarks demonstrate that our method maintains accuracy while generating significantly more concise explanations, effectively teaching models to "think when needed."

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2025) studied this question.

synapsesocial.com/papers/690e8b6ca5b062d7a4e734echttps://doi.org/10.48550/arxiv.2504.03234
Ask AI
Helpful
Bookmark
Share
View Full Paper