AIHOT 于 2026-08-18 收录了“GRPO 超越英语:多语言与非英语环境下的大规模研究”这一公开动态。以下先呈现从来源页面抓取的正文,再给出 AIHOT 摘要与 TopoReduce 编辑解读。
PUBLIC SOURCE CONTENT
已抓取公开正文公开原文内容
GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings - Apple Machine Learning Research
research area Speech and Natural Language Processing
content type paperpublished August 2026
GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
AuthorsKonstantin Dobler†**, Federico Scozzafava, Jonathan Janke, Mohamed Ali, Simon Lehnerer
View publication
Copy Bibtex
Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong crosslingual transfer: training in one language often improves performance in many others. However, specific trends are highly model- and language-dependent. In some cases, training in a particular language induces severe regressions on out-of-domain capabilities in other languages. Our analysis shows that RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.
- † Hasso Plattner Institute & ELLIS Unit Potsdam
- ** Work done while at Apple
Related readings and updates.
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
April 2, 2026research area Methods and Algorithms, research area Speech and Natural Language Processing
Despite their sophisticated general-purpose capabilities, Large Language Models (LLMs) often fail to align with diverse individual preferences because standard post-training methods, like Reinforcement Learning with Human Feedback (RLHF), optimize for a single, global objective. While Group Relative Policy Optimization (GRPO) is a widely adopted on-policy reinforcement learning framework, its group-based normalization implicitly assumes that all…
Read more
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
May 16, 2025research area Speech and Natural Language Processingconference ACL
Current Large Language Models (LLMs) are predominantly designed with English as the primary language, and even the few that are multilingual tend to exhibit strong English-centric biases. Much like speakers who might produce awkward expressions when learning a second language, LLMs often generate unnatural outputs in non-English languages, reflecting English-centric patterns in both vocabulary and grammar. Despite the importance of this issue,…
Read more
Discover opportunities in Machine Learning.
Our research in machine learning breaks new ground every day.
Work with us
content type paperpublished August 2026
GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
AuthorsKonstantin Dobler†**, Federico Scozzafava, Jonathan Janke, Mohamed Ali, Simon Lehnerer
View publication
Copy Bibtex
Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong crosslingual transfer: training in one language often improves performance in many others. However, specific trends are highly model- and language-dependent. In some cases, training in a particular language induces severe regressions on out-of-domain capabilities in other languages. Our analysis shows that RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.
- † Hasso Plattner Institute & ELLIS Unit Potsdam
- ** Work done while at Apple
Related readings and updates.
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
April 2, 2026research area Methods and Algorithms, research area Speech and Natural Language Processing
Despite their sophisticated general-purpose capabilities, Large Language Models (LLMs) often fail to align with diverse individual preferences because standard post-training methods, like Reinforcement Learning with Human Feedback (RLHF), optimize for a single, global objective. While Group Relative Policy Optimization (GRPO) is a widely adopted on-policy reinforcement learning framework, its group-based normalization implicitly assumes that all…
Read more
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
May 16, 2025research area Speech and Natural Language Processingconference ACL
Current Large Language Models (LLMs) are predominantly designed with English as the primary language, and even the few that are multilingual tend to exhibit strong English-centric biases. Much like speakers who might produce awkward expressions when learning a second language, LLMs often generate unnatural outputs in non-English languages, reflecting English-centric patterns in both vocabulary and grammar. Despite the importance of this issue,…
Read more
Discover opportunities in Machine Learning.
Our research in machine learning breaks new ground every day.
Work with us
AIHOT 摘要
一项大规模实证研究考察了 GRPO 在多语言和非英语环境下的表现,覆盖多种基础模型、训练语言及推理语言奖励设置。研究发现,以母语进行推理训练与英语推理训练之间的性能差距很小,表明 RLVR 在非英语场景下同样有效。该研究为多语言推理模型的强化学习训练提供了重要参考。
为什么值得关注
大规模多语言实验显示,非英语母语推理与英语推理的差距很小,可影响多语言模型后训练时选择训练语言的成本判断。
工程化解读
从 TopoReduce 的工程视角看,这条信息属于“论文与研究”主题。它的价值不只在于一个新产品或新观点本身,还在于说明 AI 系统正在如何影响模型接入、智能体协作、研发流程、基础设施和团队决策。实际采用前,应结合原文确认版本、适用范围、价格和运行条件。
- 发布时间:2026-08-18;AIHOT 分类:论文与研究。
- AIHOT 标签:
- AIHOT 判断:大规模多语言实验显示,非英语母语推理与英语推理的差距很小,可影响多语言模型后训练时选择训练语言的成本判断。
- AIHOT 评分:50;评分用于站内排序,不等同于独立评测结论。
TopoReduce 编辑观察
当 AI 动态进入真实生产环境,团队需要同时关注能力边界、数据来源、调用成本、权限控制和可回滚性。把单条新闻放回完整工程链路中阅读,比只看标题更有助于判断它是否适合自己的产品和工作流。