PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

Diversity-Incentivized Exploration for Versatile Reasoning

View Full Paper
ZHZ.-W. HuSZShilin ZhangYLYingwu LI

Key Points

  • DIVER framework improves reasoning by leveraging global diversity incentives for exploration in LLMs.
  • Empirical research indicates a strong correlation between global diversity and improved reasoning capacity.
  • The potential-based reward shaping mechanism enhances policy invariance while mitigating reward hacking risks.
  • DIVER demonstrates superior performance compared to existing reinforcement learning baselines across various task evaluations.

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a crucial paradigm for incentivizing reasoning capabilities in Large Language Models (LLMs). Due to vast state-action spaces and reward sparsity in reasoning tasks, existing methods often struggle with deficient exploration and poor sample efficiency. In the paper, we propose DIVER (Diversity-Incentivized Exploration for VersatilE Reasoning), an innovative framework that highlights the pivotal role of global sequence-level diversity to incentivize deep exploration for versatile reasoning. We first conduct a primary empirical study to reveal a strong positive correlation between global diversity and reasoning capacity. Building on this insight, we introduce global diversity incentives as an intrinsic reward to promote deep exploration in a semantically structured space. Incorporating the intrinsic reward, we develop a potential-based reward shaping mechanism to preserve optimal policy invariance and design simple heuristics to mitigate possible reward hacking. Experimental results show that DIVER outperforms competitive RLVR baselines with various exploration strategies on both in-domain and out-of-domain tasks, excelling in both Pass@1 and Pass@k evaluations. Our code is available at https: //github. com/NJU-RL/DIVER.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hu et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcd68d54a28a75cf1f60https://doi.org/10.48550/arxiv.2509.26209
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Outcome-based Exploration for LLM Reasoning2025
  2. 2Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration2025
  3. 3Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models2025
  4. 4Navigate the Unknown: Enhancing LLM Reasoning with Intrinsic Motivation Guided Exploration2025
  5. 5Assessing RLVR’s Efficacy in Solving Previously Intractable Problems with LLMs2026