MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Alshammari, Shaden, Wen, Kevin, Zainal, Abrar, Hamilton, Mark, Safaei, Navid, Albarakati, Sultan, Freeman, William T., Torralba, Antonio
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908980929888256
author Alshammari, Shaden
Wen, Kevin
Zainal, Abrar
Hamilton, Mark
Safaei, Navid
Albarakati, Sultan
Freeman, William T.
Torralba, Antonio
author_facet Alshammari, Shaden
Wen, Kevin
Zainal, Abrar
Hamilton, Mark
Safaei, Navid
Albarakati, Sultan
Freeman, William T.
Torralba, Antonio
contents Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce MathNet, a high-quality, large-scale, multimodal, and multilingual dataset of Olympiad-level math problems together with a benchmark for evaluating mathematical reasoning in generative models and mathematical retrieval in embedding-based systems. MathNet spans 47 countries, 17 languages, and two decades of competitions, comprising 30,676 expert-authored problems with solutions across diverse domains. In addition to the core dataset, we construct a retrieval benchmark consisting of mathematically equivalent and structurally similar problem pairs curated by human experts. MathNet supports three tasks: (i) Problem Solving, (ii) Math-Aware Retrieval, and (iii) Retrieval-Augmented Problem Solving. Experimental results show that even state-of-the-art reasoning models (78.4% for Gemini-3.1-Pro and 69.3% for GPT-5) remain challenged, while embedding models struggle to retrieve equivalent problems. We further show that retrieval-augmented generation performance is highly sensitive to retrieval quality; for example, DeepSeek-V3.2-Speciale achieves gains of up to 12%, obtaining the highest scores on the benchmark. MathNet provides the largest high-quality Olympiad dataset together with the first benchmark for evaluating mathematical problem retrieval, and we publicly release both the dataset and benchmark at https://mathnet.mit.edu.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18584
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
Alshammari, Shaden
Wen, Kevin
Zainal, Abrar
Hamilton, Mark
Safaei, Navid
Albarakati, Sultan
Freeman, William T.
Torralba, Antonio
Artificial Intelligence
Digital Libraries
Information Retrieval
Machine Learning
Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce MathNet, a high-quality, large-scale, multimodal, and multilingual dataset of Olympiad-level math problems together with a benchmark for evaluating mathematical reasoning in generative models and mathematical retrieval in embedding-based systems. MathNet spans 47 countries, 17 languages, and two decades of competitions, comprising 30,676 expert-authored problems with solutions across diverse domains. In addition to the core dataset, we construct a retrieval benchmark consisting of mathematically equivalent and structurally similar problem pairs curated by human experts. MathNet supports three tasks: (i) Problem Solving, (ii) Math-Aware Retrieval, and (iii) Retrieval-Augmented Problem Solving. Experimental results show that even state-of-the-art reasoning models (78.4% for Gemini-3.1-Pro and 69.3% for GPT-5) remain challenged, while embedding models struggle to retrieve equivalent problems. We further show that retrieval-augmented generation performance is highly sensitive to retrieval quality; for example, DeepSeek-V3.2-Speciale achieves gains of up to 12%, obtaining the highest scores on the benchmark. MathNet provides the largest high-quality Olympiad dataset together with the first benchmark for evaluating mathematical problem retrieval, and we publicly release both the dataset and benchmark at https://mathnet.mit.edu.
title MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
topic Artificial Intelligence
Digital Libraries
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2604.18584