MMATH: A Multilingual Benchmark for Mathematical Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Luo, Wenyang, Zhao, Wayne Xin, Sha, Jing, Wang, Shijin, Wen, Ji-Rong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908378911997952
author Luo, Wenyang
Zhao, Wayne Xin
Sha, Jing
Wang, Shijin
Wen, Ji-Rong
author_facet Luo, Wenyang
Zhao, Wayne Xin
Sha, Jing
Wang, Shijin
Wen, Ji-Rong
contents The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex reasoning remain underexplored, with existing efforts largely focused on simpler tasks like MGSM. To address this gap, we introduce MMATH, a benchmark for multilingual complex reasoning spanning 374 high-quality math problems across 10 typologically diverse languages. Using MMATH, we observe that even advanced models like DeepSeek R1 exhibit substantial performance disparities across languages and suffer from a critical off-target issue-generating responses in unintended languages. To address this, we explore strategies including prompting and training, demonstrating that reasoning in English and answering in target languages can simultaneously enhance performance and preserve target-language consistency. Our findings offer new insights and practical strategies for advancing the multilingual reasoning capabilities of large language models. Our code and data could be found at https://github.com/RUCAIBox/MMATH.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19126
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMATH: A Multilingual Benchmark for Mathematical Reasoning
Luo, Wenyang
Zhao, Wayne Xin
Sha, Jing
Wang, Shijin
Wen, Ji-Rong
Computation and Language
The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex reasoning remain underexplored, with existing efforts largely focused on simpler tasks like MGSM. To address this gap, we introduce MMATH, a benchmark for multilingual complex reasoning spanning 374 high-quality math problems across 10 typologically diverse languages. Using MMATH, we observe that even advanced models like DeepSeek R1 exhibit substantial performance disparities across languages and suffer from a critical off-target issue-generating responses in unintended languages. To address this, we explore strategies including prompting and training, demonstrating that reasoning in English and answering in target languages can simultaneously enhance performance and preserve target-language consistency. Our findings offer new insights and practical strategies for advancing the multilingual reasoning capabilities of large language models. Our code and data could be found at https://github.com/RUCAIBox/MMATH.
title MMATH: A Multilingual Benchmark for Mathematical Reasoning
topic Computation and Language
url https://arxiv.org/abs/2505.19126