APE-Bench: Evaluating Automated Proof Engineering for Formal Math Libraries
Fuente:
arXiv
Salvato in:
| Autori principali: | Xin, Huajian, Li, Luming, Jin, Xiaoran, Fleuriot, Jacques, Li, Wenda |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
di: Ravi, Nikil, et al.
Pubblicazione: (2026)
di: Ravi, Nikil, et al.
Pubblicazione: (2026)
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
di: Petrov, Ivo, et al.
Pubblicazione: (2025)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
di: Cao, Jialun, et al.
Pubblicazione: (2025)
di: Cao, Jialun, et al.
Pubblicazione: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
di: Li, Xin, et al.
Pubblicazione: (2025)
di: Li, Xin, et al.
Pubblicazione: (2025)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
di: Ma, Wenjie, et al.
Pubblicazione: (2025)
di: Ma, Wenjie, et al.
Pubblicazione: (2025)
Proof-RM: A Scalable and Generalizable Reward Model for Math Proof
di: Yang, Haotong, et al.
Pubblicazione: (2026)
di: Yang, Haotong, et al.
Pubblicazione: (2026)
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
di: Wang, Ruida, et al.
Pubblicazione: (2025)
di: Wang, Ruida, et al.
Pubblicazione: (2025)
Automate Knowledge Concept Tagging on Math Questions with LLMs
di: Li, Hang, et al.
Pubblicazione: (2024)
di: Li, Hang, et al.
Pubblicazione: (2024)
MUSTARD: Mastering Uniform Synthesis of Theorem and Proof Data
di: Huang, Yinya, et al.
Pubblicazione: (2024)
di: Huang, Yinya, et al.
Pubblicazione: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
di: Liu, Hongwei, et al.
Pubblicazione: (2024)
di: Liu, Hongwei, et al.
Pubblicazione: (2024)
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems
di: Kai, Ding, et al.
Pubblicazione: (2024)
di: Kai, Ding, et al.
Pubblicazione: (2024)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
di: Opedal, Andreas, et al.
Pubblicazione: (2024)
di: Opedal, Andreas, et al.
Pubblicazione: (2024)
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators
di: Lu, Qingyu, et al.
Pubblicazione: (2024)
di: Lu, Qingyu, et al.
Pubblicazione: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
di: Zhang, Boning, et al.
Pubblicazione: (2024)
di: Zhang, Boning, et al.
Pubblicazione: (2024)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
di: Yang, Tengchao, et al.
Pubblicazione: (2025)
di: Yang, Tengchao, et al.
Pubblicazione: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
di: Xu, Wenda, et al.
Pubblicazione: (2025)
di: Xu, Wenda, et al.
Pubblicazione: (2025)
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
di: Zheng, Chuanyang, et al.
Pubblicazione: (2023)
di: Zheng, Chuanyang, et al.
Pubblicazione: (2023)
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
di: Lee, Yunseung, et al.
Pubblicazione: (2026)
di: Lee, Yunseung, et al.
Pubblicazione: (2026)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
di: Xin, Yutong, et al.
Pubblicazione: (2026)
di: Xin, Yutong, et al.
Pubblicazione: (2026)
Natural Language Translation of Formal Proofs through Informalization of Proof Steps and Recursive Summarization along Proof Structure
di: Hattori, Seiji, et al.
Pubblicazione: (2025)
di: Hattori, Seiji, et al.
Pubblicazione: (2025)
FormalAlign: Automated Alignment Evaluation for Autoformalization
di: Lu, Jianqiao, et al.
Pubblicazione: (2024)
di: Lu, Jianqiao, et al.
Pubblicazione: (2024)
APE: Active Learning-based Tooling for Finding Informative Few-shot Examples for LLM-based Entity Matching
di: Qian, Kun, et al.
Pubblicazione: (2024)
di: Qian, Kun, et al.
Pubblicazione: (2024)
BenchBench: Benchmarking Automated Benchmark Generation
di: Zheng, Yandan, et al.
Pubblicazione: (2026)
di: Zheng, Yandan, et al.
Pubblicazione: (2026)
Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning
di: Zhuang, Wenwen, et al.
Pubblicazione: (2024)
di: Zhuang, Wenwen, et al.
Pubblicazione: (2024)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
di: Balunović, Mislav, et al.
Pubblicazione: (2025)
di: Balunović, Mislav, et al.
Pubblicazione: (2025)
MathDuels: Evaluating LLMs as Problem Posers and Solvers
di: Xu, Zhiqiu, et al.
Pubblicazione: (2026)
di: Xu, Zhiqiu, et al.
Pubblicazione: (2026)
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
di: Sun, Kai, et al.
Pubblicazione: (2024)
di: Sun, Kai, et al.
Pubblicazione: (2024)
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
di: An, Shengnan, et al.
Pubblicazione: (2025)
di: An, Shengnan, et al.
Pubblicazione: (2025)
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
di: Tian, Changyuan, et al.
Pubblicazione: (2025)
di: Tian, Changyuan, et al.
Pubblicazione: (2025)
Fine-Grained Natural Language Inference Based Faithfulness Evaluation for Diverse Summarisation Tasks
di: Zhang, Huajian, et al.
Pubblicazione: (2024)
di: Zhang, Huajian, et al.
Pubblicazione: (2024)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
di: Xu, Xi, et al.
Pubblicazione: (2024)
di: Xu, Xi, et al.
Pubblicazione: (2024)
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
di: Baral, Sami, et al.
Pubblicazione: (2025)
di: Baral, Sami, et al.
Pubblicazione: (2025)
APE: Selective Fine-tuning with Acceptance Criteria for Language Model Adaptation
di: Marín, Javier
Pubblicazione: (2025)
di: Marín, Javier
Pubblicazione: (2025)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
di: Zhao, Yilun, et al.
Pubblicazione: (2023)
di: Zhao, Yilun, et al.
Pubblicazione: (2023)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
di: Wang, Lei, et al.
Pubblicazione: (2024)
di: Wang, Lei, et al.
Pubblicazione: (2024)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
Formally Verified Neurosymbolic Trajectory Learning via Tensor-based Linear Temporal Logic on Finite Traces
di: Chevallier, Mark, et al.
Pubblicazione: (2025)
di: Chevallier, Mark, et al.
Pubblicazione: (2025)
Mathfish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula
di: Lucy, Li, et al.
Pubblicazione: (2024)
di: Lucy, Li, et al.
Pubblicazione: (2024)
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
di: Li, Zhong-Zhi, et al.
Pubblicazione: (2024)
di: Li, Zhong-Zhi, et al.
Pubblicazione: (2024)
How Large Language Models Balance Internal Knowledge with User and Document Assertions
di: Li, Shuowei, et al.
Pubblicazione: (2026)
di: Li, Shuowei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
di: Ravi, Nikil, et al.
Pubblicazione: (2026) -
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
di: Petrov, Ivo, et al.
Pubblicazione: (2025) -
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
di: Cao, Jialun, et al.
Pubblicazione: (2025) -
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
di: Li, Xin, et al.
Pubblicazione: (2025) -
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
di: Ma, Wenjie, et al.
Pubblicazione: (2025)