A multi-seed empirical study of GRPO post-training of DeepSeek-R1-Distill-Qwen-1.5B under a single-GPU LoRA constraint. We compare four arms (English-only, Vietnamese-translated, English with auxiliary language-consistency reward R5, and a constant-bias ablation) across three random seeds and four math benchmarks. Three findings: (i) R5 arm achieves the only positive AIME-2024 maj@8 effect (+4.4±1.9 pp over base); (ii) vanilla English GRPO shows σ=11.3 pp seed variance on AMC-23 with consistent AIME degradation, confirming single-seed claims at this scale are unreliable; (iii) 10-language MGSM evaluation shows null cross-lingual transfer — all conditions within ±0.5 pp of base. Constant-bias ablation does not fully reproduce R5 effect
Dang Vu (Mon,) studied this question.