PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 29, 20250 citationsOpen Access

Thought-Augmented Policy Optimization: Bridging External Guidance and Internal Capabilities

View Full Paper
JWJinyang WuCLChan-Yu LiaoMFMingkuan Feng

Key Points

  • TAPO demonstrates a 99% improvement over GRPO on AIME, enhancing reasoning capabilities through external guidance.
  • Effective integration of thought patterns allows TAPO to outperform existing models in various tasks and domains.
  • The framework highlights the balance between internal exploration and external guidance without compromising output quality.
  • External guidance contributes to improved explainability and readability of inference behavior in reasoning models.

Abstract

Reinforcement learning (RL) has emerged as an effective method for training reasoning models. However, existing RL approaches typically bias the model's output distribution toward reward-maximizing paths without introducing external knowledge. This limits their exploration capacity and results in a narrower reasoning capability boundary compared to base models. To address this limitation, we propose TAPO (Thought-Augmented Policy Optimization), a novel framework that augments RL by incorporating external high-level guidance ("thought patterns"). By adaptively integrating structured thoughts during training, TAPO effectively balances model-internal exploration and external guidance exploitation. Extensive experiments show that our approach significantly outperforms GRPO by 99% on AIME, 41% on AMC, and 17% on Minerva Math. Notably, these high-level thought patterns, abstracted from only 500 prior samples, generalize effectively across various tasks and models. This highlights TAPO's potential for broader applications across multiple tasks and domains. Our further analysis reveals that introducing external guidance produces powerful reasoning models with superior explainability of inference behavior and enhanced output readability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2025) studied this question.

synapsesocial.com/papers/68da58d8c1728099cfd10f0dhttps://doi.org/10.48550/arxiv.2505.15692
Ask AI
Helpful
Bookmark
Share
View Full Paper