Reliable Fine-Grained Evaluation of Natural Language Math Proofs
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Wenjie, Cojocaru, Andrei, Kolhe, Neel, Louie, Bradley, Sharif, Robin Said, Zhang, Haihan, Zhuang, Vincent, Zaharia, Matei, Min, Sewon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Reasoning Models Can Be Effective Without Thinking
by: Ma, Wenjie, et al.
Published: (2025)
by: Ma, Wenjie, et al.
Published: (2025)
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts
by: Morrison, Jacob, et al.
Published: (2026)
by: Morrison, Jacob, et al.
Published: (2026)
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
by: Garg, Ishir, et al.
Published: (2026)
by: Garg, Ishir, et al.
Published: (2026)
SIEVE: Sample-Efficient Parametric Learning from Natural Language
by: Asawa, Parth, et al.
Published: (2026)
by: Asawa, Parth, et al.
Published: (2026)
DS SERVE: A Framework for Efficient and Scalable Neural Retrieval
by: Liu, Jinjian, et al.
Published: (2025)
by: Liu, Jinjian, et al.
Published: (2025)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
InfoSynth: Information-Guided Benchmark Synthesis for LLMs
by: Garg, Ishir, et al.
Published: (2026)
by: Garg, Ishir, et al.
Published: (2026)
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
by: Garg, Ishir, et al.
Published: (2026)
by: Garg, Ishir, et al.
Published: (2026)
Fine-Grained Complexity via Quantum Natural Proofs
by: Chen, Yanlin, et al.
Published: (2025)
by: Chen, Yanlin, et al.
Published: (2025)
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023)
by: Saad-Falcon, Jon, et al.
Published: (2023)
ProofSketcher: Hybrid LLM + Lightweight Proof Checker for Reliable Math/Logic Reasoning
by: Kommuru, Kranthi, et al.
Published: (2026)
by: Kommuru, Kranthi, et al.
Published: (2026)
Natural Language Query to Configuration for Retrieval Agents
by: Pan, Melissa Z., et al.
Published: (2026)
by: Pan, Melissa Z., et al.
Published: (2026)
FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language Models
by: Liu, Yan, et al.
Published: (2024)
by: Liu, Yan, et al.
Published: (2024)
ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data
by: Patel, Liana, et al.
Published: (2024)
by: Patel, Liana, et al.
Published: (2024)
World Model on Million-Length Video And Language With Blockwise RingAttention
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
$L^*LM$: Learning Automata from Examples using Natural Language Oracles
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024)
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024)
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
Math in the Library?
by: Henry, Robin
Published: (2004)
by: Henry, Robin
Published: (2004)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
Long Context RAG Performance of Large Language Models
by: Leng, Quinn, et al.
Published: (2024)
by: Leng, Quinn, et al.
Published: (2024)
Intellectual Property Right in Indian Entrepreneural Perspective
by: Dr. Jayashree Nagorao Kolhe
Published: (2025)
by: Dr. Jayashree Nagorao Kolhe
Published: (2025)
APE-Bench: Evaluating Automated Proof Engineering for Formal Math Libraries
by: Xin, Huajian, et al.
Published: (2025)
by: Xin, Huajian, et al.
Published: (2025)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
by: Patel, Liana, et al.
Published: (2025)
by: Patel, Liana, et al.
Published: (2025)
ScenicNL: Generating Probabilistic Scenario Programs from Natural Language
by: Elmaaroufi, Karim, et al.
Published: (2024)
by: Elmaaroufi, Karim, et al.
Published: (2024)
Tooth-Diffusion: Guided 3D CBCT Synthesis with Fine-Grained Tooth Conditioning
by: Said, Said Djafar, et al.
Published: (2025)
by: Said, Said Djafar, et al.
Published: (2025)
PIMEX: Psychological Integrative Model of Existential Patterns
by: Cojocaru, Gabriel
Published: (2025)
by: Cojocaru, Gabriel
Published: (2025)
Around Context-Free Grammars -- a Normal Form, a Representation Theorem, and a Regular Approximation
by: Cojocaru, Liliana
Published: (2015)
by: Cojocaru, Liliana
Published: (2015)
On Some Complexity Results for Even Linear Languages
by: Cojocaru, Liliana
Published: (2024)
by: Cojocaru, Liliana
Published: (2024)
Proof-RM: A Scalable and Generalizable Reward Model for Math Proof
by: Yang, Haotong, et al.
Published: (2026)
by: Yang, Haotong, et al.
Published: (2026)
WARP: An Efficient Engine for Multi-Vector Retrieval
by: Scheerer, Jan Luca, et al.
Published: (2025)
by: Scheerer, Jan Luca, et al.
Published: (2025)
LEANN: A Low-Storage Vector Index
by: Wang, Yichuan, et al.
Published: (2025)
by: Wang, Yichuan, et al.
Published: (2025)
Efficiently Verifiable Proofs of Data Attribution
by: Karchmer, Ari, et al.
Published: (2025)
by: Karchmer, Ari, et al.
Published: (2025)
ElasticTok: Adaptive Tokenization for Image and Video
by: Yan, Wilson, et al.
Published: (2024)
by: Yan, Wilson, et al.
Published: (2024)
Drowning in Documents: Consequences of Scaling Reranker Inference
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)
by: Chen, Lingjiao, et al.
Published: (2026)
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
EMO: Pretraining Mixture of Experts for Emergent Modularity
by: Wang, Ryan, et al.
Published: (2026)
by: Wang, Ryan, et al.
Published: (2026)
EMVA Terminal: An Autonomous Multispectral Screening and Triage System Utilizing High-Resolution Diffuse Reflectance Spectroscopy and Artificial Intelligence
by: Cojocaru, Cristian Gabriel
Published: (2026)
by: Cojocaru, Cristian Gabriel
Published: (2026)
Similar Items
-
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026) -
Reasoning Models Can Be Effective Without Thinking
by: Ma, Wenjie, et al.
Published: (2025) -
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts
by: Morrison, Jacob, et al.
Published: (2026) -
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
by: Garg, Ishir, et al.
Published: (2026) -
SIEVE: Sample-Efficient Parametric Learning from Natural Language
by: Asawa, Parth, et al.
Published: (2026)