PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20250 citationsOpen Access

Learning to Reason without External Rewards

View Full Paper
XZXuandong ZhaoZKZhewei KangAFAosong Feng

Key Points

  • Intuitor achieves similar performance to GRPO on benchmarks, demonstrating effective reasoning without external rewards.
  • Using self-certainty as a reward signal allows for unsupervised learning across various tasks without domain-specific supervision.
  • Experiments show Intuitor can generalize better to tasks like code generation, enhancing its versatility compared to traditional methods.
  • The findings support using intrinsic signals for developing scalable AI systems, especially when external rewards are scarce.

Abstract

Training large language models (LLMs) for complex reasoning via Reinforcement Learning with Verifiable Rewards (RLVR) is effective but limited by reliance on costly, domain-specific supervision. We explore Reinforcement Learning from Internal Feedback (RLIF), a framework that enables LLMs to learn from intrinsic signals without external rewards or labeled data. We propose Intuitor, an RLIF method that uses a model's own confidence, termed self-certainty, as its sole reward signal. Intuitor replaces external rewards in Group Relative Policy Optimization (GRPO) with self-certainty scores, enabling fully unsupervised learning. Experiments demonstrate that Intuitor matches GRPO's performance on mathematical benchmarks while achieving superior generalization to out-of-domain tasks like code generation, without requiring gold solutions or test cases. Our findings show that intrinsic model signals can drive effective learning across domains, offering a scalable alternative to RLVR for autonomous AI systems where verifiable rewards are unavailable. Code is available at https://github.com/sunblaze-ucb/Intuitor

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhao et al. (2025) studied this question.

synapsesocial.com/papers/68f12bfb2107091eab27a3e0https://doi.org/10.48550/arxiv.2505.19590
Ask AI
Helpful
Bookmark
Share
View Full Paper