Factorial experiment reveals divergent training dynamics without capability gains in language models, indicating structural limitations in binary reinforcement learning.
Reinforcement learning from verifiable rewards (RLVR) has become a popular way to specialize language models for code generation, driven by a training signal as simple as whether generated code passes or fails its unit tests. Most published results, though, train a single model and report a single before/after delta, which tangles together three distinct factors: parameter count, prior code-specific pretraining, and instruction-tuning. To pull these apart, we apply one identical REINFORCE-plus-plus recipe to six model configurations built from three base checkpoints - a 3 billion parameter general model, a 3 billion parameter code-pretrained model, and a 7 billion parameter general model - each taken in both its base and its instruction-tuned release. In none of the six arms does greedy pass-at-1 on HumanEval shift significantly in either direction, and at low k the paired-bootstrap confidence intervals on pass-at-k rule out a reliable effect for every arm. Two significant high-k effects did surface for the single 7 billion parameter pair under our first sampling seed, but neither held up under an independent second seed; once replicated, no arm shows a capability effect distinguishable from zero at either seed. What does replicate across all six arms is a matter of training dynamics rather than capability: base and instruction-tuned models differ sharply in how much of their training reward comes from positive versus negative advantage under the same binary reward. The base variants finish training with an exponential moving average baseline reward well below their own pre-training diagnostic pass rate - so they learn almost entirely from negative advantage - whereas the instruction-tuned variants finish at or above that diagnostic. We offer this divergence, the still-open question of what causes it, and a set of methodological findings - among them a documented policy-gradient failure mode under sparse binary rewards and a catalog of scoring bugs that, left unfixed, would quietly produce a believable but incorrect headline result - as the primary contributions of this work.
No takes yet. Share an insight, caveat, or question.
Prapty Rahman (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: