HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ospanov, Azim, Feng, Zijin, Sun, Jiacheng, Bai, Haoli, Shen, Xin, Farnia, Farzan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BEAVER: An Efficient Deterministic LLM Verifier
by: Suresh, Tarun, et al.
Published: (2025)
by: Suresh, Tarun, et al.
Published: (2025)
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
by: Liu, Yanming, et al.
Published: (2026)
by: Liu, Yanming, et al.
Published: (2026)
Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception
by: Yang, Yunhao, et al.
Published: (2023)
by: Yang, Yunhao, et al.
Published: (2023)
Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance
by: Holzer, Nikolaus, et al.
Published: (2025)
by: Holzer, Nikolaus, et al.
Published: (2025)
LLMs as Probabilistic Minimally Adequate Teachers for DFA Learning
by: Chen, Lekai, et al.
Published: (2024)
by: Chen, Lekai, et al.
Published: (2024)
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
by: Wang, Ruida, et al.
Published: (2024)
by: Wang, Ruida, et al.
Published: (2024)
RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
Foundation Models for Logistics: Toward Certifiable, Conversational Planning Interfaces
by: Yang, Yunhao, et al.
Published: (2025)
by: Yang, Yunhao, et al.
Published: (2025)
Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs
by: Tomlinson, Kiran, et al.
Published: (2026)
by: Tomlinson, Kiran, et al.
Published: (2026)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
by: Schnabel, Tobias, et al.
Published: (2025)
by: Schnabel, Tobias, et al.
Published: (2025)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
by: Shelat, Shlok, et al.
Published: (2026)
by: Shelat, Shlok, et al.
Published: (2026)
Probabilistic Modeling of Spiking Neural Networks with Contract-Based Verification
by: Yao, Zhen, et al.
Published: (2025)
by: Yao, Zhen, et al.
Published: (2025)
Stochastic Directly-Follows Process Discovery Using Grammatical Inference
by: Alkhammash, Hanan, et al.
Published: (2023)
by: Alkhammash, Hanan, et al.
Published: (2023)
Inference of Deterministic Finite Automata via Q-Learning
by: Hosseinkhani, Elaheh, et al.
Published: (2025)
by: Hosseinkhani, Elaheh, et al.
Published: (2025)
Congruence-based Learning of Probabilistic Deterministic Finite Automata
by: Carrasco, Matías, et al.
Published: (2024)
by: Carrasco, Matías, et al.
Published: (2024)
Large Language Models and the Extended Church-Turing Thesis
by: Wiedermann, Jiří, et al.
Published: (2024)
by: Wiedermann, Jiří, et al.
Published: (2024)
Finding path and cycle counting formulae in graphs with Deep Reinforcement Learning
by: Piquenot, Jason, et al.
Published: (2024)
by: Piquenot, Jason, et al.
Published: (2024)
Are Agents Probabilistic Automata? A Trace-Based, Memory-Constrained Theory of Agentic AI
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
Computing the Reachability Value of Posterior-Deterministic POMDPs
by: Fijalkow, Nathanaël, et al.
Published: (2026)
by: Fijalkow, Nathanaël, et al.
Published: (2026)
On Synthesis of Timed Regular Expressions
by: Wang, Ziran, et al.
Published: (2025)
by: Wang, Ziran, et al.
Published: (2025)
In System Alignments we Trust! Explainable Alignments via Projections
by: Sommers, Dominique, et al.
Published: (2025)
by: Sommers, Dominique, et al.
Published: (2025)
Logic-Gated Time-Shared Feedforward Networks for Alternating Finite Automata: Exact Simulation and Learnability
by: Dhayalkar, Sahil Rajesh
Published: (2026)
by: Dhayalkar, Sahil Rajesh
Published: (2026)
Hilbert: Recursively Building Formal Proofs with Informal Reasoning
by: Varambally, Sumanth, et al.
Published: (2025)
by: Varambally, Sumanth, et al.
Published: (2025)
Closure Properties of General Grammars -- Formally Verified
by: Dvorak, Martin, et al.
Published: (2023)
by: Dvorak, Martin, et al.
Published: (2023)
Verifying Parameterized Networks Specified by Vertex-Replacement Graph Grammars
by: Iosif, Radu, et al.
Published: (2025)
by: Iosif, Radu, et al.
Published: (2025)
Probabilistic Regular Tree Priors for Scientific Symbolic Reasoning
by: Schneider, Tim, et al.
Published: (2023)
by: Schneider, Tim, et al.
Published: (2023)
Mathematical Approach in Automata and Automata Association
by: Maciel, Sergio Henrique
Published: (2020)
by: Maciel, Sergio Henrique
Published: (2020)
Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
InductionBench: LLMs Fail in the Simplest Complexity Class
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
Consistent Autoformalization for Constructing Mathematical Libraries
by: Zhang, Lan, et al.
Published: (2024)
by: Zhang, Lan, et al.
Published: (2024)
The 4/$δ$ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee
by: Dantas, PIerre, et al.
Published: (2025)
by: Dantas, PIerre, et al.
Published: (2025)
Simultaneous Task Allocation and Planning for Multi-Robots under Hierarchical Temporal Logic Specifications
by: Luo, Xusheng, et al.
Published: (2024)
by: Luo, Xusheng, et al.
Published: (2024)
Active Reward Machine Inference From Raw State Trajectories
by: Shehab, Mohamad Louai, et al.
Published: (2026)
by: Shehab, Mohamad Louai, et al.
Published: (2026)
Adaptive Bi-Level Multi-Robot Task Allocation and Learning under Uncertainty with Temporal Logic Constraints
by: Lin, Xiaoshan, et al.
Published: (2025)
by: Lin, Xiaoshan, et al.
Published: (2025)
Behavior Trees vs Executable Ontologies: a Comparative Analysis of Robot Control Paradigms
by: Boldachev, Alexander
Published: (2025)
by: Boldachev, Alexander
Published: (2025)
Joint Verification and Refinement of Language Models for Safety-Constrained Planning
by: Yang, Yunhao, et al.
Published: (2024)
by: Yang, Yunhao, et al.
Published: (2024)
SELP: Generating Safe and Efficient Task Plans for Robot Agents with Large Language Models
by: Wu, Yi, et al.
Published: (2024)
by: Wu, Yi, et al.
Published: (2024)
Integrating Supertag Features into Neural Discontinuous Constituent Parsing
by: Mielczarek, Lukas
Published: (2024)
by: Mielczarek, Lukas
Published: (2024)
Monadic Context Engineering
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Optimal Control of Logically Constrained Partially Observable and Multi-Agent Markov Decision Processes
by: Kalagarla, Krishna C., et al.
Published: (2023)
by: Kalagarla, Krishna C., et al.
Published: (2023)
Similar Items
-
BEAVER: An Efficient Deterministic LLM Verifier
by: Suresh, Tarun, et al.
Published: (2025) -
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
by: Liu, Yanming, et al.
Published: (2026) -
Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception
by: Yang, Yunhao, et al.
Published: (2023) -
Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance
by: Holzer, Nikolaus, et al.
Published: (2025) -
LLMs as Probabilistic Minimally Adequate Teachers for DFA Learning
by: Chen, Lekai, et al.
Published: (2024)