Construction-Verification: A Benchmark for Applied Mathematics in Lean 4

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Bowen, Yuan, Yi, Li, Chenyi, Wang, Ziyu, Li, Liangqi, Zhang, Bo, Li, Zhe, Wen, Zaiwen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912866928427008
author Yang, Bowen
Yuan, Yi
Li, Chenyi
Wang, Ziyu
Li, Liangqi
Zhang, Bo
Li, Zhe
Wen, Zaiwen
author_facet Yang, Bowen
Yuan, Yi
Li, Chenyi
Wang, Ziyu
Li, Liangqi
Zhang, Bo
Li, Zhe
Wen, Zaiwen
contents Recent advances in large language models have demonstrated impressive capabilities in mathematical formalization. However, existing benchmarks focus on logical verification of declarative propositions, often neglecting the task of explicitly synthesizing solutions. This limitation is particularly acute in applied mathematics domains, where the goal is frequently to derive concrete values or executable algorithms rather than solely proving theorems. To address this, we introduce a Lean 4 framework that enforces a construction-verification workflow, compelling the agent to define explicit solutions before proving their correctness. We curate a comprehensive benchmark AMBER (Applied Mathematics BEnchmark for Reasoning) spanning core domains of applied mathematics, including convex analysis, optimization, numerical algebra, and high-dimensional probability. Aside from theorem proving, our benchmark features complex tasks such as evaluation, algorithm design, and representation transformation. Experiments reveal that current models face significant difficulties with these constructive tasks. Notably, we observe that general-purpose reasoning models consistently outperform specialized theorem provers. We attribute this to a degradation of instruction following capabilities in specialized models. Fine-tuning on proof corpora appears to induce ``tactical overfitting", compromising the ability to adhere to complex constructive requirements, whereas general models retain the versatility needed for multi-task formal reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01291
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Construction-Verification: A Benchmark for Applied Mathematics in Lean 4
Yang, Bowen
Yuan, Yi
Li, Chenyi
Wang, Ziyu
Li, Liangqi
Zhang, Bo
Li, Zhe
Wen, Zaiwen
Logic in Computer Science
Recent advances in large language models have demonstrated impressive capabilities in mathematical formalization. However, existing benchmarks focus on logical verification of declarative propositions, often neglecting the task of explicitly synthesizing solutions. This limitation is particularly acute in applied mathematics domains, where the goal is frequently to derive concrete values or executable algorithms rather than solely proving theorems. To address this, we introduce a Lean 4 framework that enforces a construction-verification workflow, compelling the agent to define explicit solutions before proving their correctness. We curate a comprehensive benchmark AMBER (Applied Mathematics BEnchmark for Reasoning) spanning core domains of applied mathematics, including convex analysis, optimization, numerical algebra, and high-dimensional probability. Aside from theorem proving, our benchmark features complex tasks such as evaluation, algorithm design, and representation transformation. Experiments reveal that current models face significant difficulties with these constructive tasks. Notably, we observe that general-purpose reasoning models consistently outperform specialized theorem provers. We attribute this to a degradation of instruction following capabilities in specialized models. Fine-tuning on proof corpora appears to induce ``tactical overfitting", compromising the ability to adhere to complex constructive requirements, whereas general models retain the versatility needed for multi-task formal reasoning.
title Construction-Verification: A Benchmark for Applied Mathematics in Lean 4
topic Logic in Computer Science
url https://arxiv.org/abs/2602.01291