IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Biyani, Param, Kirtania, Shashank, Bajpai, Yasharth, Gulwani, Sumit, Tiwari, Ashish
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908877409222656
author Biyani, Param
Kirtania, Shashank
Bajpai, Yasharth
Gulwani, Sumit
Tiwari, Ashish
author_facet Biyani, Param
Kirtania, Shashank
Bajpai, Yasharth
Gulwani, Sumit
Tiwari, Ashish
contents Reliable autoformalization remains challenging even in the era of large language models (LLMs). The scarcity of high-quality training data is a major bottleneck. Expert annotation requires substantial time and deep expertise in both mathematics and theorem proving. We introduce IndiMathBench, a human-verified benchmark designed to evaluate mathematical theorem proving, curated using an AI-powered human-assisted pipeline for formalizing natural language problems in Lean. IndiMathBench is composed of 312 formal Lean 4 theorems paired with their corresponding informal problem statements, sourced from Indian Mathematics Olympiads. Through category-based retrieval, iterative compiler feedback, and multi-model ensembles, our pipeline generates candidate formalizations that experts efficiently validate via an interactive dashboard with automated quality summaries. Evaluation across multiple frontier models demonstrates that autoformalization remains challenging, with substantial gaps between syntactic validity and semantic correctness, while theorem proving success rates remain low even with iterative refinement, demonstrating that \benchmark~presents a challenging testbed for mathematical reasoning. IndiMathBench is available at https://github.com/prmbiy/IndiMathBench.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00997
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
Biyani, Param
Kirtania, Shashank
Bajpai, Yasharth
Gulwani, Sumit
Tiwari, Ashish
Artificial Intelligence
Reliable autoformalization remains challenging even in the era of large language models (LLMs). The scarcity of high-quality training data is a major bottleneck. Expert annotation requires substantial time and deep expertise in both mathematics and theorem proving. We introduce IndiMathBench, a human-verified benchmark designed to evaluate mathematical theorem proving, curated using an AI-powered human-assisted pipeline for formalizing natural language problems in Lean. IndiMathBench is composed of 312 formal Lean 4 theorems paired with their corresponding informal problem statements, sourced from Indian Mathematics Olympiads. Through category-based retrieval, iterative compiler feedback, and multi-model ensembles, our pipeline generates candidate formalizations that experts efficiently validate via an interactive dashboard with automated quality summaries. Evaluation across multiple frontier models demonstrates that autoformalization remains challenging, with substantial gaps between syntactic validity and semantic correctness, while theorem proving success rates remain low even with iterative refinement, demonstrating that \benchmark~presents a challenging testbed for mathematical reasoning. IndiMathBench is available at https://github.com/prmbiy/IndiMathBench.
title IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
topic Artificial Intelligence
url https://arxiv.org/abs/2512.00997