RLMEval: Evaluating Research-Level Neural Theorem Proving
Fuente:
arXiv
Saved in:
| Main Authors: | Poiroux, Auguste, Bosselut, Antoine, Kunčak, Viktor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024)
by: Poiroux, Auguste, et al.
Published: (2024)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
by: Kim, Kyuhee, et al.
Published: (2026)
by: Kim, Kyuhee, et al.
Published: (2026)
miniCTX: Neural Theorem Proving with (Long-)Contexts
by: Hu, Jiewen, et al.
Published: (2024)
by: Hu, Jiewen, et al.
Published: (2024)
Aristotle: IMO-level Automated Theorem Proving
by: Achim, Tudor, et al.
Published: (2025)
by: Achim, Tudor, et al.
Published: (2025)
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
PhysProver: Advancing Automatic Theorem Proving for Physics
by: Zhang, Hanning, et al.
Published: (2026)
by: Zhang, Hanning, et al.
Published: (2026)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
by: Quan, Xin, et al.
Published: (2025)
by: Quan, Xin, et al.
Published: (2025)
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
by: Chen, Luoxin, et al.
Published: (2025)
by: Chen, Luoxin, et al.
Published: (2025)
OProver: A Unified Framework for Agentic Formal Theorem Proving
by: Ma, David, et al.
Published: (2026)
by: Ma, David, et al.
Published: (2026)
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
by: Zhang, Ziyin, et al.
Published: (2025)
by: Zhang, Ziyin, et al.
Published: (2025)
Complex Reasoning over Logical Queries on Commonsense Knowledge Graphs
by: Fang, Tianqing, et al.
Published: (2024)
by: Fang, Tianqing, et al.
Published: (2024)
PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
by: Mamooler, Sepideh, et al.
Published: (2024)
by: Mamooler, Sepideh, et al.
Published: (2024)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
by: Bayazit, Deniz, et al.
Published: (2025)
by: Bayazit, Deniz, et al.
Published: (2025)
Learning to Reason with Insight for Informal Theorem Proving
by: Li, Yunhe, et al.
Published: (2026)
by: Li, Yunhe, et al.
Published: (2026)
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
by: Liu, Chengwu, et al.
Published: (2026)
by: Liu, Chengwu, et al.
Published: (2026)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
by: Gao, Silin, et al.
Published: (2025)
by: Gao, Silin, et al.
Published: (2025)
EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
by: Li, Mukai, et al.
Published: (2025)
by: Li, Mukai, et al.
Published: (2025)
Creativity in AI: Progresses and Challenges
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
by: Feng, Yiyang, et al.
Published: (2026)
by: Feng, Yiyang, et al.
Published: (2026)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
LLMs Are In-Context Bandit Reinforcement Learners
by: Monea, Giovanni, et al.
Published: (2024)
by: Monea, Giovanni, et al.
Published: (2024)
QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
by: LM-Provers, et al.
Published: (2026)
by: LM-Provers, et al.
Published: (2026)
HunyuanProver: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Neural Theorem Proving: Generating and Structuring Proofs for Formal Verification
by: Rao, Balaji, et al.
Published: (2025)
by: Rao, Balaji, et al.
Published: (2025)
MA-LoT: Model-Collaboration Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
InternLM2.5-StepProver: Advancing Automated Theorem Proving via Critic-Guided Search
by: Wu, Zijian, et al.
Published: (2024)
by: Wu, Zijian, et al.
Published: (2024)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
An In-Context Learning Agent for Formal Theorem-Proving
by: Thakur, Amitayush, et al.
Published: (2023)
by: Thakur, Amitayush, et al.
Published: (2023)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
by: Zhao, Xueliang, et al.
Published: (2024)
by: Zhao, Xueliang, et al.
Published: (2024)
Evaluating Morphological Compositional Generalization in Large Language Models
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
Evaluating Language Model Agency through Negotiations
by: Davidson, Tim R., et al.
Published: (2024)
by: Davidson, Tim R., et al.
Published: (2024)
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
by: Cao, Chuxue, et al.
Published: (2025)
by: Cao, Chuxue, et al.
Published: (2025)
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
FVEL: Interactive Formal Verification Environment with Large Language Models via Theorem Proving
by: Lin, Xiaohan, et al.
Published: (2024)
by: Lin, Xiaohan, et al.
Published: (2024)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
by: Ismayilzada, Mete, et al.
Published: (2026)
by: Ismayilzada, Mete, et al.
Published: (2026)
Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving
by: Rajaee, Sara, et al.
Published: (2025)
by: Rajaee, Sara, et al.
Published: (2025)
Similar Items
-
Reliable Evaluation and Benchmarks for Statement Autoformalization
by: Poiroux, Auguste, et al.
Published: (2024) -
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
by: Kim, Kyuhee, et al.
Published: (2026) -
miniCTX: Neural Theorem Proving with (Long-)Contexts
by: Hu, Jiewen, et al.
Published: (2024) -
Aristotle: IMO-level Automated Theorem Proving
by: Achim, Tudor, et al.
Published: (2025) -
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
by: Zheng, Chuanyang, et al.
Published: (2023)