Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Bourigault, Pauline, Ji, Xiaotong, Zimmer, Matthieu, Tutunov, Rasul, Ammar, Haitham Bou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
by: Nguyen, Tu, et al.
Published: (2026)
by: Nguyen, Tu, et al.
Published: (2026)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
by: Shen, Ziju, et al.
Published: (2025)
by: Shen, Ziju, et al.
Published: (2025)
Herald: A Natural Language Annotated Lean 4 Dataset
by: Gao, Guoxiong, et al.
Published: (2024)
by: Gao, Guoxiong, et al.
Published: (2024)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
Llemma: An Open Language Model For Mathematics
by: Azerbayev, Zhangir, et al.
Published: (2023)
by: Azerbayev, Zhangir, et al.
Published: (2023)
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
by: Cao, Chuxue, et al.
Published: (2025)
by: Cao, Chuxue, et al.
Published: (2025)
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
by: Noël, Valentin
Published: (2026)
by: Noël, Valentin
Published: (2026)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
ASP-Bench: From Natural Language to Logic Programs
by: Szeider, Stefan
Published: (2026)
by: Szeider, Stefan
Published: (2026)
Generics and Default Reasoning in Large Language Models
by: Kirkpatrick, James Ravi, et al.
Published: (2025)
by: Kirkpatrick, James Ravi, et al.
Published: (2025)
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
by: Liu, Chengwu, et al.
Published: (2026)
by: Liu, Chengwu, et al.
Published: (2026)
Towards Logically Sound Natural Language Reasoning with Logic-Enhanced Language Model Agents
by: Mensfelt, Agnieszka, et al.
Published: (2024)
by: Mensfelt, Agnieszka, et al.
Published: (2024)
DECIDER: A Dual-System Rule-Controllable Decoding Framework for Language Generation
by: Xu, Chen, et al.
Published: (2024)
by: Xu, Chen, et al.
Published: (2024)
Logic-Parametric Neuro-Symbolic NLI: Controlling Logical Formalisms for Verifiable LLM Reasoning
by: Farjami, Ali, et al.
Published: (2026)
by: Farjami, Ali, et al.
Published: (2026)
Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing
by: Świechowski, Maciej, et al.
Published: (2026)
by: Świechowski, Maciej, et al.
Published: (2026)
Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation
by: Bao, Qiming, et al.
Published: (2022)
by: Bao, Qiming, et al.
Published: (2022)
Combining Textual and Structural Information for Premise Selection in Lean
by: Petrovčič, Job, et al.
Published: (2025)
by: Petrovčič, Job, et al.
Published: (2025)
Why Can Large Language Models Generate Correct Chain-of-Thoughts?
by: Tutunov, Rasul, et al.
Published: (2023)
by: Tutunov, Rasul, et al.
Published: (2023)
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems
by: Rahman, A M Muntasir, et al.
Published: (2024)
by: Rahman, A M Muntasir, et al.
Published: (2024)
APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning
by: Ospanov, Azim, et al.
Published: (2025)
by: Ospanov, Azim, et al.
Published: (2025)
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)
by: Zimmer, Matthieu, et al.
Published: (2024)
EXPLORER: Exploration-guided Reasoning for Textual Reinforcement Learning
by: Basu, Kinjal, et al.
Published: (2024)
by: Basu, Kinjal, et al.
Published: (2024)
LLM-ARC: Enhancing LLMs with an Automated Reasoning Critic
by: Kalyanpur, Aditya, et al.
Published: (2024)
by: Kalyanpur, Aditya, et al.
Published: (2024)
Sound and Complete Neurosymbolic Reasoning with LLM-Grounded Interpretations
by: Allen, Bradley P., et al.
Published: (2025)
by: Allen, Bradley P., et al.
Published: (2025)
Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning
by: Zhang, Xinglang, et al.
Published: (2026)
by: Zhang, Xinglang, et al.
Published: (2026)
Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability
by: Zhang, Leizhen, et al.
Published: (2026)
by: Zhang, Leizhen, et al.
Published: (2026)
A Neurosymbolic Approach to Natural Language Formalization and Verification
by: Bayless, Sam, et al.
Published: (2025)
by: Bayless, Sam, et al.
Published: (2025)
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
by: Zhou, Yujun, et al.
Published: (2025)
by: Zhou, Yujun, et al.
Published: (2025)
Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
by: Li, Fangjun, et al.
Published: (2024)
by: Li, Fangjun, et al.
Published: (2024)
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
by: Lu, Jialin, et al.
Published: (2026)
by: Lu, Jialin, et al.
Published: (2026)
Lean Meets Theoretical Computer Science: Scalable Synthesis of Theorem Proving Challenges in Formal-Informal Pairs
by: Zhang, Terry Jingchen, et al.
Published: (2025)
by: Zhang, Terry Jingchen, et al.
Published: (2025)
Ontology for Policing: Conceptual Knowledge Learning for Semantic Understanding and Reasoning in Law Enforcement Reports
by: Srbinovska, Anita, et al.
Published: (2026)
by: Srbinovska, Anita, et al.
Published: (2026)
LeanTutor: Towards a Verified AI Mathematical Proof Tutor
by: Patel, Manooshree, et al.
Published: (2025)
by: Patel, Manooshree, et al.
Published: (2025)
A Reliable Common-Sense Reasoning Socialbot Built Using LLMs and Goal-Directed ASP
by: Zeng, Yankai, et al.
Published: (2024)
by: Zeng, Yankai, et al.
Published: (2024)
Defining implication relation for classical logic
by: Fu, Li
Published: (2013)
by: Fu, Li
Published: (2013)
Bridging LLMs and Symbolic Reasoning in Educational QA Systems: Insights from the XAI Challenge at IJCNN 2025
by: Nguyen, Long S. T., et al.
Published: (2025)
by: Nguyen, Long S. T., et al.
Published: (2025)
Mathematics with large language models as provers and verifiers
by: Duc, Hieu Le, et al.
Published: (2025)
by: Duc, Hieu Le, et al.
Published: (2025)
Similar Items
-
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
by: Ji, Xiaotong, et al.
Published: (2026) -
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
by: Ji, Xiaotong, et al.
Published: (2026) -
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
by: Nguyen, Tu, et al.
Published: (2026) -
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
by: Roy, Amartya, et al.
Published: (2026) -
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
by: Shen, Ziju, et al.
Published: (2025)