QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | LM-Provers, Qu, Yuxiao, Setlur, Amrith, Dekoninck, Jasper, Beeching, Edward, Li, Jia, Wu, Ian, Tunstall, Lewis, Kumar, Aviral |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
von: Qu, Yuxiao, et al.
Veröffentlicht: (2026)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2026)
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
von: Wu, Ian, et al.
Veröffentlicht: (2026)
von: Wu, Ian, et al.
Veröffentlicht: (2026)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
von: Yang, Matthew Y. R., et al.
Veröffentlicht: (2026)
von: Yang, Matthew Y. R., et al.
Veröffentlicht: (2026)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
von: Qu, Yuxiao, et al.
Veröffentlicht: (2024)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2024)
e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
CaRT: Teaching LLM Agents to Know When They Know Enough
von: Liu, Grace, et al.
Veröffentlicht: (2025)
von: Liu, Grace, et al.
Veröffentlicht: (2025)
Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
von: Setlur, Amrith, et al.
Veröffentlicht: (2026)
von: Setlur, Amrith, et al.
Veröffentlicht: (2026)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
von: Liu, Chengwu, et al.
Veröffentlicht: (2026)
von: Liu, Chengwu, et al.
Veröffentlicht: (2026)
A Unified Approach to Routing and Cascading for LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Learning from Saturated Data: Signals Beyond Correctness for LLM Training
von: Hiss, Hanno, et al.
Veröffentlicht: (2026)
von: Hiss, Hanno, et al.
Veröffentlicht: (2026)
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
von: Mündler, Niels, et al.
Veröffentlicht: (2025)
von: Mündler, Niels, et al.
Veröffentlicht: (2025)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
Benchmarking Testing in Automated Theorem Proving
von: Kim, Jongyoon, et al.
Veröffentlicht: (2026)
von: Kim, Jongyoon, et al.
Veröffentlicht: (2026)
ConStat: Performance-Based Contamination Detection in Large Language Models
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
von: Niklaus, Joel, et al.
Veröffentlicht: (2026)
von: Niklaus, Joel, et al.
Veröffentlicht: (2026)
Aristotle: IMO-level Automated Theorem Proving
von: Achim, Tudor, et al.
Veröffentlicht: (2025)
von: Achim, Tudor, et al.
Veröffentlicht: (2025)
Automated Theorem Proving for Prolog Verification
von: Mesnard, Fred, et al.
Veröffentlicht: (2026)
von: Mesnard, Fred, et al.
Veröffentlicht: (2026)
Controlled Text Generation via Language Model Arithmetic
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2023)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2023)
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
von: Petrov, Ivo, et al.
Veröffentlicht: (2026)
von: Petrov, Ivo, et al.
Veröffentlicht: (2026)
Steering LLMs for Formal Theorem Proving
von: Kirtania, Shashank, et al.
Veröffentlicht: (2025)
von: Kirtania, Shashank, et al.
Veröffentlicht: (2025)
Automated Discovery of Tactic Libraries for Interactive Theorem Proving
von: Xin, Yutong, et al.
Veröffentlicht: (2025)
von: Xin, Yutong, et al.
Veröffentlicht: (2025)
What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?
von: Kang, Katie, et al.
Veröffentlicht: (2024)
von: Kang, Katie, et al.
Veröffentlicht: (2024)
Learning to Reason with Insight for Informal Theorem Proving
von: Li, Yunhe, et al.
Veröffentlicht: (2026)
von: Li, Yunhe, et al.
Veröffentlicht: (2026)
A Theorem-Proving-Based Evaluation of Neural Semantic Parsing
von: Funakura, Hayate, et al.
Veröffentlicht: (2025)
von: Funakura, Hayate, et al.
Veröffentlicht: (2025)
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
von: Chen, Luoxin, et al.
Veröffentlicht: (2025)
von: Chen, Luoxin, et al.
Veröffentlicht: (2025)
RLMEval: Evaluating Research-Level Neural Theorem Proving
von: Poiroux, Auguste, et al.
Veröffentlicht: (2025)
von: Poiroux, Auguste, et al.
Veröffentlicht: (2025)
Lyra: Orchestrating Dual Correction in Automated Theorem Proving
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2023)
von: Zheng, Chuanyang, et al.
Veröffentlicht: (2023)
PhysProver: Advancing Automatic Theorem Proving for Physics
von: Zhang, Hanning, et al.
Veröffentlicht: (2026)
von: Zhang, Hanning, et al.
Veröffentlicht: (2026)
Lower Bounds for Public-Private Learning under Distribution Shift
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
Mechanic: Sorrifier-Driven Formal Decomposition Workflow for Automated Theorem Proving
von: Qiu, Ruichen, et al.
Veröffentlicht: (2026)
von: Qiu, Ruichen, et al.
Veröffentlicht: (2026)
miniCTX: Neural Theorem Proving with (Long-)Contexts
von: Hu, Jiewen, et al.
Veröffentlicht: (2024)
von: Hu, Jiewen, et al.
Veröffentlicht: (2024)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
von: Quan, Xin, et al.
Veröffentlicht: (2025)
von: Quan, Xin, et al.
Veröffentlicht: (2025)
OProver: A Unified Framework for Agentic Formal Theorem Proving
von: Ma, David, et al.
Veröffentlicht: (2026)
von: Ma, David, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025) -
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
von: Qu, Yuxiao, et al.
Veröffentlicht: (2026) -
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
von: Wu, Ian, et al.
Veröffentlicht: (2026) -
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
von: Petrov, Ivo, et al.
Veröffentlicht: (2025) -
Scaling Test-Time Compute Without Verification or RL is Suboptimal
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)