MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sobhani, Mahbub E, Sayeedi, Md. Faiyaz Abdullah, Mohiuddin, Tasnim, Islam, Md Mofijul, Shatabda, Swakkhar
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918303378702336
author Sobhani, Mahbub E
Sayeedi, Md. Faiyaz Abdullah
Mohiuddin, Tasnim
Islam, Md Mofijul
Shatabda, Swakkhar
author_facet Sobhani, Mahbub E
Sayeedi, Md. Faiyaz Abdullah
Mohiuddin, Tasnim
Islam, Md Mofijul
Shatabda, Swakkhar
contents Mathematical reasoning remains one of the most challenging domains for large language models (LLMs), requiring not only linguistic understanding but also structured logical deduction and numerical precision. While recent LLMs demonstrate strong general-purpose reasoning abilities, their mathematical competence across diverse languages remains underexplored. Existing benchmarks primarily focus on English or a narrow subset of high-resource languages, leaving significant gaps in assessing multilingual and cross-lingual mathematical reasoning. To address this, we introduce MATHMIST, a parallel multilingual benchmark for mathematical problem solving and reasoning. MATHMIST encompasses 2,890 parallel Bangla-English gold standard artifacts, totaling approximately 30K aligned question--answer pairs across thirteen languages, representing an extensive coverage of high-, medium-, and low-resource linguistic settings. The dataset captures linguistic variety, multiple types of problem settings, and solution synthesizing capabilities. We systematically evaluate a diverse suite of models, including open-source small and medium LLMs, proprietary systems, and multilingual-reasoning-focused models under zero-shot, chain-of-thought (CoT), perturbated reasoning, and code-switched reasoning paradigms. Our results reveal persistent deficiencies in LLMs' ability to perform consistent and interpretable mathematical reasoning across languages, with pronounced degradation in low-resource settings. All the codes and data are available at GitHub: https://github.com/mahbubhimel/MathMist
format Preprint
id arxiv_https___arxiv_org_abs_2510_14305
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
Sobhani, Mahbub E
Sayeedi, Md. Faiyaz Abdullah
Mohiuddin, Tasnim
Islam, Md Mofijul
Shatabda, Swakkhar
Computation and Language
Mathematical reasoning remains one of the most challenging domains for large language models (LLMs), requiring not only linguistic understanding but also structured logical deduction and numerical precision. While recent LLMs demonstrate strong general-purpose reasoning abilities, their mathematical competence across diverse languages remains underexplored. Existing benchmarks primarily focus on English or a narrow subset of high-resource languages, leaving significant gaps in assessing multilingual and cross-lingual mathematical reasoning. To address this, we introduce MATHMIST, a parallel multilingual benchmark for mathematical problem solving and reasoning. MATHMIST encompasses 2,890 parallel Bangla-English gold standard artifacts, totaling approximately 30K aligned question--answer pairs across thirteen languages, representing an extensive coverage of high-, medium-, and low-resource linguistic settings. The dataset captures linguistic variety, multiple types of problem settings, and solution synthesizing capabilities. We systematically evaluate a diverse suite of models, including open-source small and medium LLMs, proprietary systems, and multilingual-reasoning-focused models under zero-shot, chain-of-thought (CoT), perturbated reasoning, and code-switched reasoning paradigms. Our results reveal persistent deficiencies in LLMs' ability to perform consistent and interpretable mathematical reasoning across languages, with pronounced degradation in low-resource settings. All the codes and data are available at GitHub: https://github.com/mahbubhimel/MathMist
title MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
topic Computation and Language
url https://arxiv.org/abs/2510.14305