Automated Theorem Provers Help Improve Large Language Model Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | McGinness, Lachlan, Baumgartner, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoGalactica: A Scientific Large Language Model in Geoscience
by: Lin, Zhouhan, et al.
Published: (2023)
by: Lin, Zhouhan, et al.
Published: (2023)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
by: Ghandi, Taraneh, et al.
Published: (2026)
by: Ghandi, Taraneh, et al.
Published: (2026)
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
by: Li, Shenghao
Published: (2025)
by: Li, Shenghao
Published: (2025)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
by: Chua, Jaymari, et al.
Published: (2025)
by: Chua, Jaymari, et al.
Published: (2025)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
by: Poddar, Aheli, et al.
Published: (2025)
by: Poddar, Aheli, et al.
Published: (2025)
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
by: Raman, Vishal, et al.
Published: (2025)
by: Raman, Vishal, et al.
Published: (2025)
Transforming Computer Security and Public Trust Through the Exploration of Fine-Tuning Large Language Models
by: Crumrine, Garrett, et al.
Published: (2024)
by: Crumrine, Garrett, et al.
Published: (2024)
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
by: Ansell, Rebecca, et al.
Published: (2026)
by: Ansell, Rebecca, et al.
Published: (2026)
Controlled Territory and Conflict Tracking (CONTACT): (Geo-)Mapping Occupied Territory from Open Source Intelligence
by: Mandal, Paul K., et al.
Published: (2025)
by: Mandal, Paul K., et al.
Published: (2025)
Logic.py: Bridging the Gap between LLMs and Constraint Solvers
by: Kesseli, Pascal, et al.
Published: (2025)
by: Kesseli, Pascal, et al.
Published: (2025)
Critical Insights into Leading Conversational AI Models
by: Kohli, Urja, et al.
Published: (2025)
by: Kohli, Urja, et al.
Published: (2025)
REMoH: A Reflective Evolution of Multi-objective Heuristics approach via Large Language Models
by: Forniés-Tabuenca, Diego, et al.
Published: (2025)
by: Forniés-Tabuenca, Diego, et al.
Published: (2025)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
by: Zmanovskii, Nikita
Published: (2025)
by: Zmanovskii, Nikita
Published: (2025)
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
by: Loi, Dario, et al.
Published: (2025)
by: Loi, Dario, et al.
Published: (2025)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
by: Zhang, Peiyan, et al.
Published: (2025)
by: Zhang, Peiyan, et al.
Published: (2025)
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
by: Xu, Qiyuan, et al.
Published: (2026)
by: Xu, Qiyuan, et al.
Published: (2026)
ChatGPT4PCG Competition: Character-like Level Generation for Science Birds
by: Taveekitworachai, Pittawat, et al.
Published: (2023)
by: Taveekitworachai, Pittawat, et al.
Published: (2023)
Open-TI: Open Traffic Intelligence with Augmented Language Model
by: Da, Longchao, et al.
Published: (2023)
by: Da, Longchao, et al.
Published: (2023)
Reinforced Language Models for Sequential Decision Making
by: Dilkes, Jim, et al.
Published: (2025)
by: Dilkes, Jim, et al.
Published: (2025)
HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
by: Jang, Yoonjin, et al.
Published: (2026)
by: Jang, Yoonjin, et al.
Published: (2026)
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
by: Wu, Qiming, et al.
Published: (2024)
by: Wu, Qiming, et al.
Published: (2024)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
by: Steele, Brady, et al.
Published: (2026)
by: Steele, Brady, et al.
Published: (2026)
IntSat: Integer Linear Programming by Conflict-Driven Constraint-Learning
by: Nieuwenhuis, Robert, et al.
Published: (2024)
by: Nieuwenhuis, Robert, et al.
Published: (2024)
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
by: Cui, Jian, et al.
Published: (2026)
by: Cui, Jian, et al.
Published: (2026)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
ChatGPT4PCG 2 Competition: Prompt Engineering for Science Birds Level Generation
by: Taveekitworachai, Pittawat, et al.
Published: (2024)
by: Taveekitworachai, Pittawat, et al.
Published: (2024)
Dynamics of COVID-19 Misinformation: An Analysis of Conspiracy Theories, Fake Remedies, and False Reports
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
CAPE: Corrective Actions from Precondition Errors using Large Language Models
by: Raman, Shreyas Sundara, et al.
Published: (2022)
by: Raman, Shreyas Sundara, et al.
Published: (2022)
An Explainable Collaborative Dialogue System using a Theory of Mind
by: Cohen, Philip R., et al.
Published: (2023)
by: Cohen, Philip R., et al.
Published: (2023)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations
by: Karlsson, Robin, et al.
Published: (2026)
by: Karlsson, Robin, et al.
Published: (2026)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
by: Hossain, Ariyan, et al.
Published: (2025)
by: Hossain, Ariyan, et al.
Published: (2025)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
CoupleEvo: Evolving Heuristics for Coupled Optimization Problems Using Large Language Models
by: Bömer, Thomas, et al.
Published: (2026)
by: Bömer, Thomas, et al.
Published: (2026)
Who Leads in the Shadows? ERGM and Centrality Analysis of Congressional Democrats on Bluesky
by: Hew, Gordon, et al.
Published: (2025)
by: Hew, Gordon, et al.
Published: (2025)
Cross-Subreddit Behavior as Open-Source Indicators of Coordinated Influence: A Case Study of r/Sino & r/China
by: Pilaud, Manon, et al.
Published: (2025)
by: Pilaud, Manon, et al.
Published: (2025)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
by: Kim, Kyuhee, et al.
Published: (2026)
by: Kim, Kyuhee, et al.
Published: (2026)
Can AI Assist in Olympiad Coding
by: Ren, Samuel
Published: (2025)
by: Ren, Samuel
Published: (2025)
Game of Thought: Robust Information Seeking with Large Language Models Using Game Theory
by: Cui, Langyuan, et al.
Published: (2026)
by: Cui, Langyuan, et al.
Published: (2026)
Similar Items
-
GeoGalactica: A Scientific Large Language Model in Geoscience
by: Lin, Zhouhan, et al.
Published: (2023) -
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
by: Ghandi, Taraneh, et al.
Published: (2026) -
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
by: Li, Shenghao
Published: (2025) -
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
by: Chua, Jaymari, et al.
Published: (2025) -
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
by: Poddar, Aheli, et al.
Published: (2025)