Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Yebo, Liu, Zixiang, Li, Yaoming, Yang, Zhizhuo, Xu, Xinye, Ye, Bowen, Yuan, Weijun, Wang, Zihan, Yang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
by: Dekoninck, Jasper, et al.
Published: (2025)
by: Dekoninck, Jasper, et al.
Published: (2025)
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
by: Ye, Bingyang, et al.
Published: (2026)
by: Ye, Bingyang, et al.
Published: (2026)
MUSTARD: Mastering Uniform Synthesis of Theorem and Proof Data
by: Huang, Yinya, et al.
Published: (2024)
by: Huang, Yinya, et al.
Published: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
by: Tian, Yuxuan, et al.
Published: (2025)
by: Tian, Yuxuan, et al.
Published: (2025)
Autograding Mathematical Induction Proofs with Natural Language Processing
by: Zhao, Chenyan, et al.
Published: (2024)
by: Zhao, Chenyan, et al.
Published: (2024)
ProofOptimizer: Training Language Models to Simplify Proofs without Human Demonstrations
by: Gu, Alex, et al.
Published: (2025)
by: Gu, Alex, et al.
Published: (2025)
Do We Need Frontier Models to Verify Mathematical Proofs?
by: Naik, Aaditya, et al.
Published: (2026)
by: Naik, Aaditya, et al.
Published: (2026)
Hierarchical Attention Generates Better Proofs
by: Chen, Jianlong, et al.
Published: (2025)
by: Chen, Jianlong, et al.
Published: (2025)
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
by: Thakur, Amitayush, et al.
Published: (2025)
by: Thakur, Amitayush, et al.
Published: (2025)
HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement
by: Hu, Jilin, et al.
Published: (2025)
by: Hu, Jilin, et al.
Published: (2025)
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
by: Yuan, Aomufei, et al.
Published: (2025)
by: Yuan, Aomufei, et al.
Published: (2025)
The Proof is in the Almond Cookies
by: van Trijp, Remi, et al.
Published: (2025)
by: van Trijp, Remi, et al.
Published: (2025)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
by: Wu, Qinzhuo, et al.
Published: (2026)
by: Wu, Qinzhuo, et al.
Published: (2026)
SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
by: Xin, Yue, et al.
Published: (2025)
by: Xin, Yue, et al.
Published: (2025)
Neuro-Symbolic Integration Brings Causal and Reliable Reasoning Proofs
by: Yang, Sen, et al.
Published: (2023)
by: Yang, Sen, et al.
Published: (2023)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
by: Ravi, Nikil, et al.
Published: (2026)
by: Ravi, Nikil, et al.
Published: (2026)
Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming
by: Chakraborty, Saikat, et al.
Published: (2024)
by: Chakraborty, Saikat, et al.
Published: (2024)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages
by: Chen, Zui, et al.
Published: (2025)
by: Chen, Zui, et al.
Published: (2025)
LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches
by: He, Linyang, et al.
Published: (2026)
by: He, Linyang, et al.
Published: (2026)
A Survey on Human-Centric LLMs
by: Wang, Jing Yi, et al.
Published: (2024)
by: Wang, Jing Yi, et al.
Published: (2024)
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
by: Schmitt, Johannes, et al.
Published: (2025)
by: Schmitt, Johannes, et al.
Published: (2025)
Translating Informal Proofs into Formal Proofs Using a Chain of States
by: Wang, Ziyu, et al.
Published: (2025)
by: Wang, Ziyu, et al.
Published: (2025)
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
by: Dorn, Diego, et al.
Published: (2024)
by: Dorn, Diego, et al.
Published: (2024)
QED: An Open-Source Multi-Agent System for Generating Mathematical Proofs on Open Problems
by: An, Chenyang, et al.
Published: (2026)
by: An, Chenyang, et al.
Published: (2026)
Proof-RM: A Scalable and Generalizable Reward Model for Math Proof
by: Yang, Haotong, et al.
Published: (2026)
by: Yang, Haotong, et al.
Published: (2026)
Hilbert: Recursively Building Formal Proofs with Informal Reasoning
by: Varambally, Sumanth, et al.
Published: (2025)
by: Varambally, Sumanth, et al.
Published: (2025)
AutoVerus: Automated Proof Generation for Rust Code
by: Yang, Chenyuan, et al.
Published: (2024)
by: Yang, Chenyuan, et al.
Published: (2024)
A Primer in Post-Training Reasoning Data: What We Know About How It Works
by: Li, Yaoming, et al.
Published: (2026)
by: Li, Yaoming, et al.
Published: (2026)
Aletheia tackles FirstProof autonomously
by: Feng, Tony, et al.
Published: (2026)
by: Feng, Tony, et al.
Published: (2026)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
Prove Your Point!: Bringing Proof-Enhancement Principles to Argumentative Essay Generation
by: Xiao, Ruiyu, et al.
Published: (2024)
by: Xiao, Ruiyu, et al.
Published: (2024)
Automatic Generation of Question Hints for Mathematics Problems using Large Language Models in Educational Technology
by: Tonga, Junior Cedric, et al.
Published: (2024)
by: Tonga, Junior Cedric, et al.
Published: (2024)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
by: Ma, Wenjie, et al.
Published: (2025)
by: Ma, Wenjie, et al.
Published: (2025)
A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants
by: Bayazıt, Barış, et al.
Published: (2025)
by: Bayazıt, Barış, et al.
Published: (2025)
Solving Inequality Proofs with Large Language Models
by: Lu, Pan, et al.
Published: (2025)
by: Lu, Pan, et al.
Published: (2025)
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems
by: Rahman, A M Muntasir, et al.
Published: (2024)
by: Rahman, A M Muntasir, et al.
Published: (2024)
SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
by: Guo, Beichen, et al.
Published: (2025)
by: Guo, Beichen, et al.
Published: (2025)
GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing
by: Cong, Peizhuang, et al.
Published: (2024)
by: Cong, Peizhuang, et al.
Published: (2024)
Similar Items
-
The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
by: Dekoninck, Jasper, et al.
Published: (2025) -
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
by: Ye, Bingyang, et al.
Published: (2026) -
MUSTARD: Mastering Uniform Synthesis of Theorem and Proof Data
by: Huang, Yinya, et al.
Published: (2024) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
by: Tian, Yuxuan, et al.
Published: (2025) -
Autograding Mathematical Induction Proofs with Natural Language Processing
by: Zhao, Chenyan, et al.
Published: (2024)