MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Balunović, Mislav, Dekoninck, Jasper, Petrov, Ivo, Jovanović, Nikola, Vechev, Martin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909990575407104
author Balunović, Mislav
Dekoninck, Jasper
Petrov, Ivo
Jovanović, Nikola
Vechev, Martin
author_facet Balunović, Mislav
Dekoninck, Jasper
Petrov, Ivo
Jovanović, Nikola
Vechev, Martin
contents The rapid advancement of reasoning capabilities in large language models (LLMs) has led to notable improvements on mathematical benchmarks. However, many of the most commonly used evaluation datasets (e.g., AIME 2024) are widely available online, making it difficult to disentangle genuine reasoning from potential memorization. Furthermore, these benchmarks do not evaluate proof-writing capabilities, which are crucial for many mathematical tasks. To address this, we introduce MathArena, a new benchmark based on the following key insight: recurring math competitions provide a stream of high-quality, challenging problems that can be used for real-time evaluation of LLMs. By evaluating models as soon as new problems are released, we effectively eliminate the risk of contamination. Using this framework, we find strong signs of contamination in AIME 2024. Nonetheless, evaluations on harder competitions, such as CMIMC 2025, demonstrate impressive reasoning capabilities in top-performing models. MathArena is also the first benchmark for proof-writing capabilities. On IMO 2025, top models achieve slightly less than 40%, demonstrating both notable progress and significant room for improvement. So far, we have evaluated over $50$ models across seven competitions, totaling $162$ problems. As an evolving benchmark, MathArena will continue to track the progress of LLMs on newly released competitions, ensuring rigorous and up-to-date evaluation of mathematical reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23281
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Balunović, Mislav
Dekoninck, Jasper
Petrov, Ivo
Jovanović, Nikola
Vechev, Martin
Artificial Intelligence
Computation and Language
The rapid advancement of reasoning capabilities in large language models (LLMs) has led to notable improvements on mathematical benchmarks. However, many of the most commonly used evaluation datasets (e.g., AIME 2024) are widely available online, making it difficult to disentangle genuine reasoning from potential memorization. Furthermore, these benchmarks do not evaluate proof-writing capabilities, which are crucial for many mathematical tasks. To address this, we introduce MathArena, a new benchmark based on the following key insight: recurring math competitions provide a stream of high-quality, challenging problems that can be used for real-time evaluation of LLMs. By evaluating models as soon as new problems are released, we effectively eliminate the risk of contamination. Using this framework, we find strong signs of contamination in AIME 2024. Nonetheless, evaluations on harder competitions, such as CMIMC 2025, demonstrate impressive reasoning capabilities in top-performing models. MathArena is also the first benchmark for proof-writing capabilities. On IMO 2025, top models achieve slightly less than 40%, demonstrating both notable progress and significant room for improvement. So far, we have evaluated over $50$ models across seven competitions, totaling $162$ problems. As an evolving benchmark, MathArena will continue to track the progress of LLMs on newly released competitions, ensuring rigorous and up-to-date evaluation of mathematical reasoning.
title MathArena: Evaluating LLMs on Uncontaminated Math Competitions
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.23281