APE-Bench: Evaluating Automated Proof Engineering for Formal Math Libraries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xin, Huajian, Li, Luming, Jin, Xiaoran, Fleuriot, Jacques, Li, Wenda |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
von: Ravi, Nikil, et al.
Veröffentlicht: (2026)
von: Ravi, Nikil, et al.
Veröffentlicht: (2026)
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
von: Cao, Jialun, et al.
Veröffentlicht: (2025)
von: Cao, Jialun, et al.
Veröffentlicht: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)
Proof-RM: A Scalable and Generalizable Reward Model for Math Proof
von: Yang, Haotong, et al.
Veröffentlicht: (2026)
von: Yang, Haotong, et al.
Veröffentlicht: (2026)
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
von: Wang, Ruida, et al.
Veröffentlicht: (2025)
von: Wang, Ruida, et al.
Veröffentlicht: (2025)
Automate Knowledge Concept Tagging on Math Questions with LLMs
von: Li, Hang, et al.
Veröffentlicht: (2024)
von: Li, Hang, et al.
Veröffentlicht: (2024)
MUSTARD: Mastering Uniform Synthesis of Theorem and Proof Data
von: Huang, Yinya, et al.
Veröffentlicht: (2024)
von: Huang, Yinya, et al.
Veröffentlicht: (2024)
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems
von: Kai, Ding, et al.
Veröffentlicht: (2024)
von: Kai, Ding, et al.
Veröffentlicht: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators
von: Lu, Qingyu, et al.
Veröffentlicht: (2024)
von: Lu, Qingyu, et al.
Veröffentlicht: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
von: Zhang, Boning, et al.
Veröffentlicht: (2024)
von: Zhang, Boning, et al.
Veröffentlicht: (2024)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
von: Yang, Tengchao, et al.
Veröffentlicht: (2025)
von: Yang, Tengchao, et al.
Veröffentlicht: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2023)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2023)
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
von: Lee, Yunseung, et al.
Veröffentlicht: (2026)
von: Lee, Yunseung, et al.
Veröffentlicht: (2026)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
von: Xin, Yutong, et al.
Veröffentlicht: (2026)
von: Xin, Yutong, et al.
Veröffentlicht: (2026)
Natural Language Translation of Formal Proofs through Informalization of Proof Steps and Recursive Summarization along Proof Structure
von: Hattori, Seiji, et al.
Veröffentlicht: (2025)
von: Hattori, Seiji, et al.
Veröffentlicht: (2025)
FormalAlign: Automated Alignment Evaluation for Autoformalization
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
APE: Active Learning-based Tooling for Finding Informative Few-shot Examples for LLM-based Entity Matching
von: Qian, Kun, et al.
Veröffentlicht: (2024)
von: Qian, Kun, et al.
Veröffentlicht: (2024)
BenchBench: Benchmarking Automated Benchmark Generation
von: Zheng, Yandan, et al.
Veröffentlicht: (2026)
von: Zheng, Yandan, et al.
Veröffentlicht: (2026)
Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning
von: Zhuang, Wenwen, et al.
Veröffentlicht: (2024)
von: Zhuang, Wenwen, et al.
Veröffentlicht: (2024)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
MathDuels: Evaluating LLMs as Problem Posers and Solvers
von: Xu, Zhiqiu, et al.
Veröffentlicht: (2026)
von: Xu, Zhiqiu, et al.
Veröffentlicht: (2026)
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
von: Sun, Kai, et al.
Veröffentlicht: (2024)
von: Sun, Kai, et al.
Veröffentlicht: (2024)
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
von: An, Shengnan, et al.
Veröffentlicht: (2025)
von: An, Shengnan, et al.
Veröffentlicht: (2025)
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
von: Tian, Changyuan, et al.
Veröffentlicht: (2025)
von: Tian, Changyuan, et al.
Veröffentlicht: (2025)
Fine-Grained Natural Language Inference Based Faithfulness Evaluation for Diverse Summarisation Tasks
von: Zhang, Huajian, et al.
Veröffentlicht: (2024)
von: Zhang, Huajian, et al.
Veröffentlicht: (2024)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
von: Xu, Xi, et al.
Veröffentlicht: (2024)
von: Xu, Xi, et al.
Veröffentlicht: (2024)
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
von: Baral, Sami, et al.
Veröffentlicht: (2025)
von: Baral, Sami, et al.
Veröffentlicht: (2025)
APE: Selective Fine-tuning with Acceptance Criteria for Language Model Adaptation
von: Marín, Javier
Veröffentlicht: (2025)
von: Marín, Javier
Veröffentlicht: (2025)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
Formally Verified Neurosymbolic Trajectory Learning via Tensor-based Linear Temporal Logic on Finite Traces
von: Chevallier, Mark, et al.
Veröffentlicht: (2025)
von: Chevallier, Mark, et al.
Veröffentlicht: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
Mathfish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula
von: Lucy, Li, et al.
Veröffentlicht: (2024)
von: Lucy, Li, et al.
Veröffentlicht: (2024)
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
von: Li, Zhong-Zhi, et al.
Veröffentlicht: (2024)
von: Li, Zhong-Zhi, et al.
Veröffentlicht: (2024)
How Large Language Models Balance Internal Knowledge with User and Document Assertions
von: Li, Shuowei, et al.
Veröffentlicht: (2026)
von: Li, Shuowei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
von: Ravi, Nikil, et al.
Veröffentlicht: (2026) -
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
von: Petrov, Ivo, et al.
Veröffentlicht: (2025) -
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
von: Cao, Jialun, et al.
Veröffentlicht: (2025) -
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
von: Li, Xin, et al.
Veröffentlicht: (2025) -
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)