Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Shelat, Shlok, Raval, Jay, Roy, Souvik, Gaur, Manas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance
by: Holzer, Nikolaus, et al.
Published: (2025)
by: Holzer, Nikolaus, et al.
Published: (2025)
BEAVER: An Efficient Deterministic LLM Verifier
by: Suresh, Tarun, et al.
Published: (2025)
by: Suresh, Tarun, et al.
Published: (2025)
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
by: Ospanov, Azim, et al.
Published: (2025)
by: Ospanov, Azim, et al.
Published: (2025)
Computing the Reachability Value of Posterior-Deterministic POMDPs
by: Fijalkow, Nathanaël, et al.
Published: (2026)
by: Fijalkow, Nathanaël, et al.
Published: (2026)
SELP: Generating Safe and Efficient Task Plans for Robot Agents with Large Language Models
by: Wu, Yi, et al.
Published: (2024)
by: Wu, Yi, et al.
Published: (2024)
Are Agents Probabilistic Automata? A Trace-Based, Memory-Constrained Theory of Agentic AI
by: Koohestani, Roham, et al.
Published: (2025)
by: Koohestani, Roham, et al.
Published: (2025)
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
by: Liu, Yanming, et al.
Published: (2026)
by: Liu, Yanming, et al.
Published: (2026)
Integrating Supertag Features into Neural Discontinuous Constituent Parsing
by: Mielczarek, Lukas
Published: (2024)
by: Mielczarek, Lukas
Published: (2024)
Monadic Context Engineering
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Computability of Agentic Systems
by: Viriyasuthee, Chatavut
Published: (2026)
by: Viriyasuthee, Chatavut
Published: (2026)
TuringQ: Benchmarking AI Comprehension in Theory of Computation
by: Zahraei, Pardis Sadat, et al.
Published: (2024)
by: Zahraei, Pardis Sadat, et al.
Published: (2024)
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
by: Li, Zelong, et al.
Published: (2024)
by: Li, Zelong, et al.
Published: (2024)
Simultaneous Task Allocation and Planning for Multi-Robots under Hierarchical Temporal Logic Specifications
by: Luo, Xusheng, et al.
Published: (2024)
by: Luo, Xusheng, et al.
Published: (2024)
Adaptive Bi-Level Multi-Robot Task Allocation and Learning under Uncertainty with Temporal Logic Constraints
by: Lin, Xiaoshan, et al.
Published: (2025)
by: Lin, Xiaoshan, et al.
Published: (2025)
Logic-Gated Time-Shared Feedforward Networks for Alternating Finite Automata: Exact Simulation and Learnability
by: Dhayalkar, Sahil Rajesh
Published: (2026)
by: Dhayalkar, Sahil Rajesh
Published: (2026)
Probabilistic Modeling of Spiking Neural Networks with Contract-Based Verification
by: Yao, Zhen, et al.
Published: (2025)
by: Yao, Zhen, et al.
Published: (2025)
Stochastic Directly-Follows Process Discovery Using Grammatical Inference
by: Alkhammash, Hanan, et al.
Published: (2023)
by: Alkhammash, Hanan, et al.
Published: (2023)
Inference of Deterministic Finite Automata via Q-Learning
by: Hosseinkhani, Elaheh, et al.
Published: (2025)
by: Hosseinkhani, Elaheh, et al.
Published: (2025)
LLMs as Probabilistic Minimally Adequate Teachers for DFA Learning
by: Chen, Lekai, et al.
Published: (2024)
by: Chen, Lekai, et al.
Published: (2024)
Congruence-based Learning of Probabilistic Deterministic Finite Automata
by: Carrasco, Matías, et al.
Published: (2024)
by: Carrasco, Matías, et al.
Published: (2024)
Large Language Models and the Extended Church-Turing Thesis
by: Wiedermann, Jiří, et al.
Published: (2024)
by: Wiedermann, Jiří, et al.
Published: (2024)
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
by: Wang, Ruida, et al.
Published: (2024)
by: Wang, Ruida, et al.
Published: (2024)
Finding path and cycle counting formulae in graphs with Deep Reinforcement Learning
by: Piquenot, Jason, et al.
Published: (2024)
by: Piquenot, Jason, et al.
Published: (2024)
Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception
by: Yang, Yunhao, et al.
Published: (2023)
by: Yang, Yunhao, et al.
Published: (2023)
Foundation Models for Logistics: Toward Certifiable, Conversational Planning Interfaces
by: Yang, Yunhao, et al.
Published: (2025)
by: Yang, Yunhao, et al.
Published: (2025)
On Synthesis of Timed Regular Expressions
by: Wang, Ziran, et al.
Published: (2025)
by: Wang, Ziran, et al.
Published: (2025)
In System Alignments we Trust! Explainable Alignments via Projections
by: Sommers, Dominique, et al.
Published: (2025)
by: Sommers, Dominique, et al.
Published: (2025)
What is Formal Verification without Specifications? A Survey on mining LTL Specifications
by: Neider, Daniel, et al.
Published: (2025)
by: Neider, Daniel, et al.
Published: (2025)
Fine-Tuning Language Models Using Formal Methods Feedback
by: Yang, Yunhao, et al.
Published: (2023)
by: Yang, Yunhao, et al.
Published: (2023)
RepV: Safety-Separable Latent Spaces for Scalable Neurosymbolic Plan Verification
by: Yang, Yunhao, et al.
Published: (2025)
by: Yang, Yunhao, et al.
Published: (2025)
Benchmarking Testing in Automated Theorem Proving
by: Kim, Jongyoon, et al.
Published: (2026)
by: Kim, Jongyoon, et al.
Published: (2026)
Towards Autoformalization of LLM-generated Outputs for Requirement Verification
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
A Survey of Cellular Automata: Types, Dynamics, Non-uniformity and Applications
by: Bhattacharjee, Kamalika, et al.
Published: (2016)
by: Bhattacharjee, Kamalika, et al.
Published: (2016)
Learning Probabilistic Temporal Logic Specifications for Stochastic Systems
by: Roy, Rajarshi, et al.
Published: (2025)
by: Roy, Rajarshi, et al.
Published: (2025)
On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning
by: Nowak, Franz, et al.
Published: (2024)
by: Nowak, Franz, et al.
Published: (2024)
Mining Beyond the Bools: Learning Data Transformations and Temporal Specifications
by: Kouteili, Sam Nicholas, et al.
Published: (2026)
by: Kouteili, Sam Nicholas, et al.
Published: (2026)
MASA: LLM-Driven Multi-Agent Systems for Autoformalization
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
About Time: Model-free Reinforcement Learning with Timed Reward Machines
by: Roy, Rajarshi, et al.
Published: (2025)
by: Roy, Rajarshi, et al.
Published: (2025)
Similar Items
-
RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
by: Jin, Hyundong, et al.
Published: (2025) -
Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance
by: Holzer, Nikolaus, et al.
Published: (2025) -
BEAVER: An Efficient Deterministic LLM Verifier
by: Suresh, Tarun, et al.
Published: (2025) -
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
by: Ospanov, Azim, et al.
Published: (2025) -
Computing the Reachability Value of Posterior-Deterministic POMDPs
by: Fijalkow, Nathanaël, et al.
Published: (2026)