Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gulati, Aryan, Miranda, Brando, Chen, Eric, Xia, Emily, Fronsdal, Kai, Dumont, Bruno, Obbad, Elyas, Koyejo, Sanmi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915465246277632
author Gulati, Aryan
Miranda, Brando
Chen, Eric
Xia, Emily
Fronsdal, Kai
Dumont, Bruno
Obbad, Elyas
Koyejo, Sanmi
author_facet Gulati, Aryan
Miranda, Brando
Chen, Eric
Xia, Emily
Fronsdal, Kai
Dumont, Bruno
Obbad, Elyas
Koyejo, Sanmi
contents Current mathematical reasoning benchmarks for large language models (LLMs) are approaching saturation, with some achieving > 90% accuracy, and are increasingly compromised by training-set contamination. We introduce Putnam-AXIOM, a benchmark of 522 university-level competition problems drawn from the prestigious William Lowell Putnam Mathematical Competition, and Putnam-AXIOM Variation, an unseen companion set of 100 functional variants generated by programmatically perturbing variables and constants. The variation protocol produces an unlimited stream of equally difficult, unseen instances -- yielding a contamination-resilient test bed. On the Original set, OpenAI's o1-preview -- the strongest evaluated model -- scores 41.9%, but its accuracy drops by 19.6% (46.8% relative decrease) on the paired Variations. The remaining eighteen models show the same downward trend, ten of them with non-overlapping 95% confidence intervals. These gaps suggest memorization and highlight the necessity of dynamic benchmarks. We complement "boxed" accuracy with Teacher-Forced Accuracy (TFA), a lightweight metric that directly scores reasoning traces and automates natural language proof evaluations. Putnam-AXIOM therefore provides a rigorous, contamination-resilient evaluation framework for assessing advanced mathematical reasoning of LLMs. Data and evaluation code are publicly available at https://github.com/brando90/putnam-axiom.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08292
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
Gulati, Aryan
Miranda, Brando
Chen, Eric
Xia, Emily
Fronsdal, Kai
Dumont, Bruno
Obbad, Elyas
Koyejo, Sanmi
Computation and Language
Artificial Intelligence
Machine Learning
Logic in Computer Science
Neural and Evolutionary Computing
68T20, 68T05, 68Q32
F.2.2; I.2.3; I.2.6; I.2.8
Current mathematical reasoning benchmarks for large language models (LLMs) are approaching saturation, with some achieving > 90% accuracy, and are increasingly compromised by training-set contamination. We introduce Putnam-AXIOM, a benchmark of 522 university-level competition problems drawn from the prestigious William Lowell Putnam Mathematical Competition, and Putnam-AXIOM Variation, an unseen companion set of 100 functional variants generated by programmatically perturbing variables and constants. The variation protocol produces an unlimited stream of equally difficult, unseen instances -- yielding a contamination-resilient test bed. On the Original set, OpenAI's o1-preview -- the strongest evaluated model -- scores 41.9%, but its accuracy drops by 19.6% (46.8% relative decrease) on the paired Variations. The remaining eighteen models show the same downward trend, ten of them with non-overlapping 95% confidence intervals. These gaps suggest memorization and highlight the necessity of dynamic benchmarks. We complement "boxed" accuracy with Teacher-Forced Accuracy (TFA), a lightweight metric that directly scores reasoning traces and automates natural language proof evaluations. Putnam-AXIOM therefore provides a rigorous, contamination-resilient evaluation framework for assessing advanced mathematical reasoning of LLMs. Data and evaluation code are publicly available at https://github.com/brando90/putnam-axiom.
title Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
Logic in Computer Science
Neural and Evolutionary Computing
68T20, 68T05, 68Q32
F.2.2; I.2.3; I.2.6; I.2.8
url https://arxiv.org/abs/2508.08292