MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Fuente:
arXiv
Saved in:
| Main Authors: | Balunović, Mislav, Dekoninck, Jasper, Petrov, Ivo, Jovanović, Nikola, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Large Language Models are Advanced Anonymizers
by: Staab, Robin, et al.
Published: (2024)
by: Staab, Robin, et al.
Published: (2024)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
by: Dekoninck, Jasper, et al.
Published: (2025)
by: Dekoninck, Jasper, et al.
Published: (2025)
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
by: Petrov, Ivo, et al.
Published: (2026)
by: Petrov, Ivo, et al.
Published: (2026)
Adaptive Generation of Bias-Eliciting Questions for LLMs
by: Staab, Robin, et al.
Published: (2025)
by: Staab, Robin, et al.
Published: (2025)
A Unified Approach to Routing and Cascading for LLMs
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
ToolFuzz -- Automated Agent Tool Testing
by: Milev, Ivan, et al.
Published: (2025)
by: Milev, Ivan, et al.
Published: (2025)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
CuTS: Customizable Tabular Synthetic Data Generation
by: Vero, Mark, et al.
Published: (2023)
by: Vero, Mark, et al.
Published: (2023)
Beyond Memorization: Violating Privacy Via Inference with Large Language Models
by: Staab, Robin, et al.
Published: (2023)
by: Staab, Robin, et al.
Published: (2023)
Evading Data Contamination Detection for Language Models is (too) Easy
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
by: Mitra, Arindam, et al.
Published: (2024)
by: Mitra, Arindam, et al.
Published: (2024)
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange
by: Satpute, Ankit, et al.
Published: (2024)
by: Satpute, Ankit, et al.
Published: (2024)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
by: Tian, Shi-Yu, et al.
Published: (2025)
by: Tian, Shi-Yu, et al.
Published: (2025)
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
MegaMath: Pushing the Limits of Open Math Corpora
by: Zhou, Fan, et al.
Published: (2025)
by: Zhou, Fan, et al.
Published: (2025)
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
by: An, Shengnan, et al.
Published: (2025)
by: An, Shengnan, et al.
Published: (2025)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
by: Liu, Xianyang, et al.
Published: (2025)
by: Liu, Xianyang, et al.
Published: (2025)
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
MathBuddy: A Multimodal System for Affective Math Tutoring
by: Kar, Debanjana, et al.
Published: (2025)
by: Kar, Debanjana, et al.
Published: (2025)
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward Passes
by: Christ, Bryan R., et al.
Published: (2024)
by: Christ, Bryan R., et al.
Published: (2024)
Learning from Saturated Data: Signals Beyond Correctness for LLM Training
by: Hiss, Hanno, et al.
Published: (2026)
by: Hiss, Hanno, et al.
Published: (2026)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
by: Ma, Wenjie, et al.
Published: (2025)
by: Ma, Wenjie, et al.
Published: (2025)
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
by: Li, Chengpeng, et al.
Published: (2023)
by: Li, Chengpeng, et al.
Published: (2023)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
by: Chen, Nuo, et al.
Published: (2024)
by: Chen, Nuo, et al.
Published: (2024)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
by: Li, Hang, et al.
Published: (2024)
by: Li, Hang, et al.
Published: (2024)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
by: Mahran, Mariam, et al.
Published: (2025)
by: Mahran, Mariam, et al.
Published: (2025)
Adversarial Math Word Problem Generation
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
by: Fang, Meng, et al.
Published: (2024)
by: Fang, Meng, et al.
Published: (2024)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
UTMath: Math Evaluation with Unit Test via Reasoning-to-Coding Thoughts
by: Yang, Bo, et al.
Published: (2024)
by: Yang, Bo, et al.
Published: (2024)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
by: Chen, Zui, et al.
Published: (2024)
by: Chen, Zui, et al.
Published: (2024)
Similar Items
-
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026) -
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
by: Balunović, Mislav, et al.
Published: (2025) -
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
by: Petrov, Ivo, et al.
Published: (2025) -
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025) -
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
by: Dekoninck, Jasper, et al.
Published: (2024)