PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 7, 20250 citationsOpen Access

Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning

View Full Paper
GSGuijin SonJHJiwoo HongHKHyunwoo Ko

Key Points

  • Using Qwen2.5-1.5B Math with ORM achieved a score of 35.8 on MCLM, while BF on MR1-1.5B reached 35.2.
  • BF delivers a 20-point improvement on English AIME but only a 1.94-point gain on average across other languages.
  • Test-time scaling methods like ORM and BF show performance similar to traditional scaling when limited by inference FLOPs.
  • The findings suggest that test-time scaling does not generalize effectively to multilingual mathematical tasks.

Abstract

Scaling pre-training compute has proven effective for achieving mulitlinguality, but does the same hold for test-time scaling? In this work, we introduce MCLM, a multilingual math benchmark featuring competition-level problems in 55 languages. We test three test-time scaling methods-Outcome Reward Modeling (ORM), Process Reward Modeling (ORM), and Budget Forcing (BF)-on both Qwen2.5-1.5B Math and MR1-1.5B, a multilingual LLM we trained for extended reasoning. Our experiments show that using Qwen2.5-1.5B Math with ORM achieves a score of 35.8 on MCLM, while BF on MR1-1.5B attains 35.2. Although "thinking LLMs" have recently garnered significant attention, we find that their performance is comparable to traditional scaling methods like best-of-N once constrained to similar levels of inference FLOPs. Moreover, while BF yields a 20-point improvement on English AIME, it provides only a 1.94-point average gain across other languages-a pattern consistent across the other test-time scaling methods we studied-higlighting that test-time scaling may not generalize as effectively to multilingual tasks. To foster further research, we release MCLM, MR1-1.5B, and evaluation results.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Son et al. (2025) studied this question.

synapsesocial.com/papers/68e585d0b1e78cc4e5f463e3https://doi.org/10.48550/arxiv.2502.17407
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Multilingual Test-Time Scaling via Initial Thought Transfer2025
  2. 2MathScale: Scaling Instruction Tuning for Mathematical Reasoning2024 · 2 citations
  3. 3Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning2025
  4. 4MMATH: A Multilingual Benchmark for Mathematical Reasoning2025
  5. 5Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study2026