Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rajaee, Sara, Choenni, Rochelle, Shutova, Ekaterina, Monz, Christof
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916957714907136
author Rajaee, Sara
Choenni, Rochelle
Shutova, Ekaterina
Monz, Christof
author_facet Rajaee, Sara
Choenni, Rochelle
Shutova, Ekaterina
Monz, Christof
contents While the reasoning abilities of large language models (LLMs) continue to advance, it remains unclear how such ability varies across languages in multilingual LLMs and whether different languages produce reasoning paths that complement each other. To investigate this question, we train a reward model to rank generated responses for a given question across languages. Our results show that our cross-lingual reward model substantially improves mathematical reasoning performance compared to using reward modeling within a single language, benefiting even high-resource languages. While English often exhibits the highest performance in multilingual models, we find that cross-lingual sampling particularly benefits English under low sampling budgets. Our findings reveal new opportunities to improve multilingual reasoning by leveraging the complementary strengths of diverse languages.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15811
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
Rajaee, Sara
Choenni, Rochelle
Shutova, Ekaterina
Monz, Christof
Computation and Language
Artificial Intelligence
While the reasoning abilities of large language models (LLMs) continue to advance, it remains unclear how such ability varies across languages in multilingual LLMs and whether different languages produce reasoning paths that complement each other. To investigate this question, we train a reward model to rank generated responses for a given question across languages. Our results show that our cross-lingual reward model substantially improves mathematical reasoning performance compared to using reward modeling within a single language, benefiting even high-resource languages. While English often exhibits the highest performance in multilingual models, we find that cross-lingual sampling particularly benefits English under low sampling budgets. Our findings reveal new opportunities to improve multilingual reasoning by leveraging the complementary strengths of diverse languages.
title Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.15811