This research demonstrates the structural limits of self-generated curricula in reinforcement learning, suggesting challenges in scalability.
Reinforcement learning from verifiable reward (RLVR) relies on human-curated problem sets paired with automatic evaluators. This paper asks what happens when the problem set comes from the model itself. On HumanEval with Qwen3-8B, a pipeline that generates candidate programming problems and retains only those the base model solves inconsistently produces 22 usable problems. GRPO on those 22 problems alone raises pass@1 from 63.4% to 76.8%. The same GRPO recipe on 664 curated problems from HumanEval and MBPP reaches 84.1%; self-generation recovers 65% of that gain on two orders of magnitude less training data. Three structural limits cap this approach as a route to continued self-improvement: - Iteration does not accumulate: a second GRPO round on problems calibrated against the first-round checkpoint converges back to it. - Task-type imbalance destroys transfer: a curriculum with 84% abduction-style problems drops pass@1 to 61.0%, below the pre-training baseline, by shifting the output-format prior away from what HumanEval rewards. - Inapplicability at scale: at 32B parameters the learnability window closes entirely. The base model solves essentially every problem it can formulate with a verified reference solution, and calibration has nothing to accept. A per-problem inspection of the 8B checkpoint shows that eight of ten cases where the trained adapter beats its base are indentation fixes on code the base already writes correctly. One case is a genuine reasoning gain. GRPO at this scale teaches the model to format its answers for the evaluator; new reasoning is the minority component. The same closed-loop structure that makes self-generation cost-free is what caps it.
No takes yet. Share an insight, caveat, or question.
Jesus Tabares Montilla (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: