GRPO Beyond English: Multilingual RLVR Study

This paper, by Konstantin Dobler, Federico Scozzafava, Jonathan Janke, Mohamed Ali, and Simon Lehnerer (Hasso Plattner Institute & ELLIS Unit Potsdam, with work done while at Apple), presents a large-scale empirical study of GRPO-based Reinforcement Learning with Verifiable Rewards (RLVR) in non-English and multilingual settings. The study spans multiple base models, training languages, and reasoning language reward configurations.

The main findings are that training a model to reason in the native language often leaves only a small performance gap relative to training for English reasoning, and that crosslingual transfer is strong: training in one language frequently improves performance in other languages. However, trends are highly model- and language-dependent, and in some cases training in a particular language causes severe regressions in out-of-domain capabilities in other languages. The authors conclude that extending RLVR/GRPO beyond English can produce broad crosslingual gains, but only if evaluation is broad enough to detect language-specific regressions.

GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

View Original