PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 5, 20250 citationsOpen Access

LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning

View Full Paper
WCWeizhe ChenSKSven KoenigBDBistra Dilkina

Key Points

  • Length-aware sampling improves learning effectiveness in large language models during reasoning tasks.
  • The study presents a novel meta-RLVR algorithm for dynamic sampling based on average response length.
  • Multiple base models and datasets were evaluated, with consistent improvements observed across the board.
  • An ablation study offered insights into integrating length signals, highlighting future research directions.

Abstract

Since the release of Deepseek-R1, reinforcement learning with verifiable rewards (RLVR) has become a central approach for training large language models (LLMs) on reasoning tasks. Recent work has largely focused on modifying loss functions to make RLVR more efficient and effective. In this paper, motivated by studies of overthinking in LLMs, we propose Length-aware Sampling for Policy Optimization (LSPO), a novel meta-RLVR algorithm that dynamically selects training data at each step based on the average response length. We evaluate LSPO across multiple base models and datasets, demonstrating that it consistently improves learning effectiveness. In addition, we conduct a detailed ablation study to examine alternative ways of incorporating length signals into dynamic sampling, offering further insights and highlighting promising directions for future research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2025) studied this question.

synapsesocial.com/papers/68e25382d6d66a53c2474b48https://doi.org/10.48550/arxiv.2510.01459
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning2025
  2. 2SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM2025
  3. 3Sample-efficient LLM Optimization with Reset Replay2025
  4. 4Value Augmented Sampling for Language Model Alignment and Personalization2024
  5. 5Teaching Large Language Models to Reason with Reinforcement Learning2024 · 3 citations