Multi-seed empirical study of GRPO post-training on DeepSeek-R1-Distill-Qwen-1.5B under a single-GPU LoRA (r=16) constraint. Four arms — English-only (A1), Vietnamese-translated (A2), English with auxiliary language-consistency reward R5 (A3), and a content-free constant-bias control (A4) — are compared across three seeds on AMC-23, MATH-500, AIME-2024, and 10-language MGSM. Three orthogonal findings: (i) only A3 produces a positive AIME-2024 maj@8 effect (37.8 ± 1.9%, +4.4 pp over base, robust across seeds); (ii) vanilla English GRPO shows σ = 11.3 pp seed variance on AMC-23 and degrades AIME-2024 pass@1 in 3/3 seeds (mean −12.2 pp), confirming that single-seed claims at sub-3B + LoRA scale are unreliable; (iii) on 10-language MGSM all four conditions converge to within ±0.5 pp of the untrained base — null cross-lingual transfer at this scale. The A4 constant-bias control does not reproduce the R5 effect (A3 − A4 = +5.58 pp on AIME-2024 maj@8, 95% bootstrap CI +1.13, +11.10), ruling out a pure reward-magnitude explanation.
Dang Vu (Tue,) studied this question.