Towards Robust Mathematical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luong, Thang, Hwang, Dawsen, Nguyen, Hoang H., Ghiasi, Golnaz, Chervonyi, Yuri, Seo, Insuk, Kim, Junsu, Bingham, Garrett, Lee, Jonathan, Mishra, Swaroop, Zhai, Alex, Hu, Clara Huiyi, Michalewski, Henryk, Kim, Jimin, Ahn, Jeonghyun, Bae, Junhwi, Song, Xingyou, Trinh, Trieu H., Le, Quoc V., Jung, Junehyuk
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912685873954816
author Luong, Thang
Hwang, Dawsen
Nguyen, Hoang H.
Ghiasi, Golnaz
Chervonyi, Yuri
Seo, Insuk
Kim, Junsu
Bingham, Garrett
Lee, Jonathan
Mishra, Swaroop
Zhai, Alex
Hu, Clara Huiyi
Michalewski, Henryk
Kim, Jimin
Ahn, Jeonghyun
Bae, Junhwi
Song, Xingyou
Trinh, Trieu H.
Le, Quoc V.
Jung, Junehyuk
author_facet Luong, Thang
Hwang, Dawsen
Nguyen, Hoang H.
Ghiasi, Golnaz
Chervonyi, Yuri
Seo, Insuk
Kim, Junsu
Bingham, Garrett
Lee, Jonathan
Mishra, Swaroop
Zhai, Alex
Hu, Clara Huiyi
Michalewski, Henryk
Kim, Jimin
Ahn, Jeonghyun
Bae, Junhwi
Song, Xingyou
Trinh, Trieu H.
Le, Quoc V.
Jung, Junehyuk
contents Finding the right north-star metrics is highly critical for advancing the mathematical reasoning capabilities of foundation models, especially given that existing evaluations are either too easy or only focus on getting correct short answers. To address these issues, we present IMO-Bench, a suite of advanced reasoning benchmarks, vetted by a panel of top specialists and that specifically targets the level of the International Mathematical Olympiad (IMO), the most prestigious venue for young mathematicians. IMO-AnswerBench first tests models on 400 diverse Olympiad problems with verifiable short answers. IMO-Proof Bench is the next-level evaluation for proof-writing capabilities, which includes both basic and advanced IMO level problems as well as detailed grading guidelines to facilitate automatic grading. These benchmarks played a crucial role in our historic achievement of the gold-level performance at IMO 2025 with Gemini Deep Think (Luong and Lockhart, 2025). Our model achieved 80.0% on IMO-AnswerBench and 65.7% on the advanced IMO-Proof Bench, surpassing the best non-Gemini models by large margins of 6.9% and 42.4% respectively. We also showed that autograders built with Gemini reasoning correlate well with human evaluations and construct IMO-GradingBench, with 1000 human gradings on proofs, to enable further progress in automatic evaluation of long-form answers. We hope that IMO-Bench will help the community towards advancing robust mathematical reasoning and release it at https://imobench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Robust Mathematical Reasoning
Luong, Thang
Hwang, Dawsen
Nguyen, Hoang H.
Ghiasi, Golnaz
Chervonyi, Yuri
Seo, Insuk
Kim, Junsu
Bingham, Garrett
Lee, Jonathan
Mishra, Swaroop
Zhai, Alex
Hu, Clara Huiyi
Michalewski, Henryk
Kim, Jimin
Ahn, Jeonghyun
Bae, Junhwi
Song, Xingyou
Trinh, Trieu H.
Le, Quoc V.
Jung, Junehyuk
Computation and Language
Artificial Intelligence
Finding the right north-star metrics is highly critical for advancing the mathematical reasoning capabilities of foundation models, especially given that existing evaluations are either too easy or only focus on getting correct short answers. To address these issues, we present IMO-Bench, a suite of advanced reasoning benchmarks, vetted by a panel of top specialists and that specifically targets the level of the International Mathematical Olympiad (IMO), the most prestigious venue for young mathematicians. IMO-AnswerBench first tests models on 400 diverse Olympiad problems with verifiable short answers. IMO-Proof Bench is the next-level evaluation for proof-writing capabilities, which includes both basic and advanced IMO level problems as well as detailed grading guidelines to facilitate automatic grading. These benchmarks played a crucial role in our historic achievement of the gold-level performance at IMO 2025 with Gemini Deep Think (Luong and Lockhart, 2025). Our model achieved 80.0% on IMO-AnswerBench and 65.7% on the advanced IMO-Proof Bench, surpassing the best non-Gemini models by large margins of 6.9% and 42.4% respectively. We also showed that autograders built with Gemini reasoning correlate well with human evaluations and construct IMO-GradingBench, with 1000 human gradings on proofs, to enable further progress in automatic evaluation of long-form answers. We hope that IMO-Bench will help the community towards advancing robust mathematical reasoning and release it at https://imobench.github.io/.
title Towards Robust Mathematical Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.01846