Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | McGinness, Lachlan, Baumgartner, Peter |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automated Theorem Provers Help Improve Large Language Model Reasoning
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024)
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024)
Highlighting Case Studies in LLM Literature Review of Interdisciplinary System Science
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025)
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025)
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025)
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025)
Large Language Models Imitate Logical Reasoning, but at what Cost?
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025)
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025)
The AlphaPhysics Term Rewriting System for Marking Algebraic Expressions in Physics Exams
von: Baumgartner, Peter, et al.
Veröffentlicht: (2025)
von: Baumgartner, Peter, et al.
Veröffentlicht: (2025)
CON-FOLD -- Explainable Machine Learning with Confidence
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024)
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024)
Overview of AI Grading of Physics Olympiad Exams
von: McGinness, Lachlan
Veröffentlicht: (2025)
von: McGinness, Lachlan
Veröffentlicht: (2025)
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
von: Chen, Luoxin, et al.
Veröffentlicht: (2025)
von: Chen, Luoxin, et al.
Veröffentlicht: (2025)
Can Large Language Models Correctly Interpret Equations with Errors?
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025)
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025)
EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
von: Li, Mukai, et al.
Veröffentlicht: (2025)
von: Li, Mukai, et al.
Veröffentlicht: (2025)
The Benefits and Challenges of a Quantum Computing Concept Inventory
von: McGinness, Lachlan
Veröffentlicht: (2025)
von: McGinness, Lachlan
Veröffentlicht: (2025)
PhysProver: Advancing Automatic Theorem Proving for Physics
von: Zhang, Hanning, et al.
Veröffentlicht: (2026)
von: Zhang, Hanning, et al.
Veröffentlicht: (2026)
HunyuanProver: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
InternLM2.5-StepProver: Advancing Automated Theorem Proving via Critic-Guided Search
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
von: Tsoukalas, George, et al.
Veröffentlicht: (2024)
von: Tsoukalas, George, et al.
Veröffentlicht: (2024)
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
von: Shen, Ziju, et al.
Veröffentlicht: (2025)
von: Shen, Ziju, et al.
Veröffentlicht: (2025)
An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning
von: Kaur, Navdeep, et al.
Veröffentlicht: (2025)
von: Kaur, Navdeep, et al.
Veröffentlicht: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs
von: Li, Guchan, et al.
Veröffentlicht: (2026)
von: Li, Guchan, et al.
Veröffentlicht: (2026)
Vibe Coding an LLM-powered Theorem Prover
von: Hou, Zhe
Veröffentlicht: (2026)
von: Hou, Zhe
Veröffentlicht: (2026)
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
von: Zhang, Ziyin, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2025)
LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers
von: Olausson, Theo X., et al.
Veröffentlicht: (2023)
von: Olausson, Theo X., et al.
Veröffentlicht: (2023)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
von: de Curtò, J., et al.
Veröffentlicht: (2025)
von: de Curtò, J., et al.
Veröffentlicht: (2025)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
von: Bertsch, Amanda, et al.
Veröffentlicht: (2025)
von: Bertsch, Amanda, et al.
Veröffentlicht: (2025)
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
von: Wang, Ruida, et al.
Veröffentlicht: (2025)
von: Wang, Ruida, et al.
Veröffentlicht: (2025)
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
von: Ren, Z. Z., et al.
Veröffentlicht: (2025)
von: Ren, Z. Z., et al.
Veröffentlicht: (2025)
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning
von: Wang, Jianing, et al.
Veröffentlicht: (2026)
von: Wang, Jianing, et al.
Veröffentlicht: (2026)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
von: Kim, Minwu, et al.
Veröffentlicht: (2025)
von: Kim, Minwu, et al.
Veröffentlicht: (2025)
CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models
von: Irwin, Lucas, et al.
Veröffentlicht: (2025)
von: Irwin, Lucas, et al.
Veröffentlicht: (2025)
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities
von: Hong, Zhaochen, et al.
Veröffentlicht: (2025)
von: Hong, Zhaochen, et al.
Veröffentlicht: (2025)
Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
von: Ding, Yuyang, et al.
Veröffentlicht: (2024)
von: Ding, Yuyang, et al.
Veröffentlicht: (2024)
Are LLM Belief Updates Consistent with Bayes' Theorem?
von: Imran, Sohaib, et al.
Veröffentlicht: (2025)
von: Imran, Sohaib, et al.
Veröffentlicht: (2025)
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
GraphMind: Theorem Selection and Conclusion Generation Framework with Dynamic GNN for LLM Reasoning
von: Li, Yutong, et al.
Veröffentlicht: (2025)
von: Li, Yutong, et al.
Veröffentlicht: (2025)
HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement
von: Hu, Jilin, et al.
Veröffentlicht: (2025)
von: Hu, Jilin, et al.
Veröffentlicht: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Aristotle: IMO-level Automated Theorem Proving
von: Achim, Tudor, et al.
Veröffentlicht: (2025)
von: Achim, Tudor, et al.
Veröffentlicht: (2025)
Competition-Level Problems are Effective LLM Evaluators
von: Huang, Yiming, et al.
Veröffentlicht: (2023)
von: Huang, Yiming, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Automated Theorem Provers Help Improve Large Language Model Reasoning
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024) -
Highlighting Case Studies in LLM Literature Review of Interdisciplinary System Science
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025) -
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025) -
Large Language Models Imitate Logical Reasoning, but at what Cost?
von: McGinness, Lachlan, et al.
Veröffentlicht: (2025) -
The AlphaPhysics Term Rewriting System for Marking Algebraic Expressions in Physics Exams
von: Baumgartner, Peter, et al.
Veröffentlicht: (2025)