The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Raimondi, Bianca, Pivi, Francesco, Evangelista, Davide, Gabbrielli, Maurizio
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914367502548992
author Raimondi, Bianca
Pivi, Francesco
Evangelista, Davide
Gabbrielli, Maurizio
author_facet Raimondi, Bianca
Pivi, Francesco
Evangelista, Davide
Gabbrielli, Maurizio
contents The evaluation of Large Language Models (LLMs) on mathematical reasoning has largely focused on elementary problems, competition-style questions, or formal theorem proving, leaving graduate-level and computational mathematics relatively underexplored. We introduce CompMath-MCQ, a new benchmark dataset for assessing LLMs on advanced mathematical reasoning in a multiple-choice setting. The dataset consists of 1{,}500 originally authored questions by professors of graduate-level courses, covering topics including Linear Algebra, Numerical Optimization, Vector Calculus, Probability, and Python-based scientific computing. Three option choices are provided for each question, with exactly one of them being correct. To ensure the absence of data leakage, all questions are newly created and not sourced from existing materials. The validity of questions is verified through a procedure based on cross-LLM disagreement, followed by manual expert review. By adopting a multiple-choice format, our dataset enables objective, reproducible, and bias-free evaluation through lm_eval library. Baseline results with state-of-the-art LLMs indicate that advanced computational mathematical reasoning remains a significant challenge. We release CompMath-MCQ at the following link: https://github.com/biancaraimondi/CompMath-MCQ.git
format Preprint
id arxiv_https___arxiv_org_abs_2603_03334
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
Raimondi, Bianca
Pivi, Francesco
Evangelista, Davide
Gabbrielli, Maurizio
Computation and Language
The evaluation of Large Language Models (LLMs) on mathematical reasoning has largely focused on elementary problems, competition-style questions, or formal theorem proving, leaving graduate-level and computational mathematics relatively underexplored. We introduce CompMath-MCQ, a new benchmark dataset for assessing LLMs on advanced mathematical reasoning in a multiple-choice setting. The dataset consists of 1{,}500 originally authored questions by professors of graduate-level courses, covering topics including Linear Algebra, Numerical Optimization, Vector Calculus, Probability, and Python-based scientific computing. Three option choices are provided for each question, with exactly one of them being correct. To ensure the absence of data leakage, all questions are newly created and not sourced from existing materials. The validity of questions is verified through a procedure based on cross-LLM disagreement, followed by manual expert review. By adopting a multiple-choice format, our dataset enables objective, reproducible, and bias-free evaluation through lm_eval library. Baseline results with state-of-the-art LLMs indicate that advanced computational mathematical reasoning remains a significant challenge. We release CompMath-MCQ at the following link: https://github.com/biancaraimondi/CompMath-MCQ.git
title The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
topic Computation and Language
url https://arxiv.org/abs/2603.03334