InductionBench: LLMs Fail in the Simplest Complexity Class
Fuente:
arXiv
Salvato in:
| Autori principali: | Hua, Wenyue, Wong, Tyler, Fei, Sun, Pan, Liangming, Jardine, Adam, Wang, William Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance
di: Holzer, Nikolaus, et al.
Pubblicazione: (2025)
di: Holzer, Nikolaus, et al.
Pubblicazione: (2025)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
di: Schnabel, Tobias, et al.
Pubblicazione: (2025)
di: Schnabel, Tobias, et al.
Pubblicazione: (2025)
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
di: Wang, Ruida, et al.
Pubblicazione: (2024)
di: Wang, Ruida, et al.
Pubblicazione: (2024)
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
di: Ospanov, Azim, et al.
Pubblicazione: (2025)
di: Ospanov, Azim, et al.
Pubblicazione: (2025)
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
di: Li, Zelong, et al.
Pubblicazione: (2024)
di: Li, Zelong, et al.
Pubblicazione: (2024)
Waiting Nets: State Classes and Taxonomy
di: Hélouët, Loïc, et al.
Pubblicazione: (2022)
di: Hélouët, Loïc, et al.
Pubblicazione: (2022)
LLMs as Probabilistic Minimally Adequate Teachers for DFA Learning
di: Chen, Lekai, et al.
Pubblicazione: (2024)
di: Chen, Lekai, et al.
Pubblicazione: (2024)
Computational Complexity of Alignments
di: Schwanen, Christopher T., et al.
Pubblicazione: (2026)
di: Schwanen, Christopher T., et al.
Pubblicazione: (2026)
On the Complexity of Language Membership for Probabilistic Words
di: Amarilli, Antoine, et al.
Pubblicazione: (2025)
di: Amarilli, Antoine, et al.
Pubblicazione: (2025)
Operational State Complexity of Block Languages
di: Duarte, Guilherme, et al.
Pubblicazione: (2024)
di: Duarte, Guilherme, et al.
Pubblicazione: (2024)
On the Representation and State Complexity of Block Languages
di: Duarte, Guilherme, et al.
Pubblicazione: (2024)
di: Duarte, Guilherme, et al.
Pubblicazione: (2024)
Descriptional Complexity of Finite Automata -- Selected Highlights
di: Salomaa, Arto, et al.
Pubblicazione: (2023)
di: Salomaa, Arto, et al.
Pubblicazione: (2023)
Star Complexity of Parikh Images of Languages over Infinite Alphabets
di: Danieli, Yoav
Pubblicazione: (2026)
di: Danieli, Yoav
Pubblicazione: (2026)
Fine-Grained Complexity of Ambiguity Problems on Automata and Directed Graphs
di: Drabik, Karolina, et al.
Pubblicazione: (2025)
di: Drabik, Karolina, et al.
Pubblicazione: (2025)
A Complexity Bound for Determinisation of Min-Plus Weighted Automata
di: Almagor, Shaull, et al.
Pubblicazione: (2026)
di: Almagor, Shaull, et al.
Pubblicazione: (2026)
On the Complexity of Computing the Co-lexicographic Width of a Regular Language
di: Becker, Ruben, et al.
Pubblicazione: (2024)
di: Becker, Ruben, et al.
Pubblicazione: (2024)
Unconditional Time and Space Complexity Lower Bounds for Intersection Non-Emptiness
di: Wehar, Michael
Pubblicazione: (2025)
di: Wehar, Michael
Pubblicazione: (2025)
Fine-Tuning Language Models Using Formal Methods Feedback
di: Yang, Yunhao, et al.
Pubblicazione: (2023)
di: Yang, Yunhao, et al.
Pubblicazione: (2023)
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
di: Liu, Yanming, et al.
Pubblicazione: (2026)
di: Liu, Yanming, et al.
Pubblicazione: (2026)
Finding path and cycle counting formulae in graphs with Deep Reinforcement Learning
di: Piquenot, Jason, et al.
Pubblicazione: (2024)
di: Piquenot, Jason, et al.
Pubblicazione: (2024)
Mind the Gap: A Formal Investigation of the Relationship Between Log and Model Complexity -- Extended Version
di: Schalk, Patrizia, et al.
Pubblicazione: (2025)
di: Schalk, Patrizia, et al.
Pubblicazione: (2025)
Structural Abstraction and Refinement for Probabilistic Programs
di: Li, Guanyan, et al.
Pubblicazione: (2025)
di: Li, Guanyan, et al.
Pubblicazione: (2025)
Multimodal Pretrained Models for Verifiable Sequential Decision-Making: Planning, Grounding, and Perception
di: Yang, Yunhao, et al.
Pubblicazione: (2023)
di: Yang, Yunhao, et al.
Pubblicazione: (2023)
Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs
di: Tomlinson, Kiran, et al.
Pubblicazione: (2026)
di: Tomlinson, Kiran, et al.
Pubblicazione: (2026)
Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions
di: Zhang, Lan, et al.
Pubblicazione: (2025)
di: Zhang, Lan, et al.
Pubblicazione: (2025)
Foundation Models for Logistics: Toward Certifiable, Conversational Planning Interfaces
di: Yang, Yunhao, et al.
Pubblicazione: (2025)
di: Yang, Yunhao, et al.
Pubblicazione: (2025)
The CFG Complexity of Singleton Sets
di: Fortnow, Lance, et al.
Pubblicazione: (2024)
di: Fortnow, Lance, et al.
Pubblicazione: (2024)
The rIC3 Hardware Model Checker
di: Su, Yuheng, et al.
Pubblicazione: (2025)
di: Su, Yuheng, et al.
Pubblicazione: (2025)
Vulnerabilities Analysis and Secure Controlling for Unmanned Aerial System Based on Reactive Synthesis
di: Yang, Dong, et al.
Pubblicazione: (2024)
di: Yang, Dong, et al.
Pubblicazione: (2024)
Fast and General Automatic Differentiation for Finite-State Methods
di: Yang, Lucas Ondel, et al.
Pubblicazione: (2026)
di: Yang, Lucas Ondel, et al.
Pubblicazione: (2026)
LangSAT: A Novel Framework Combining NLP and Reinforcement Learning for SAT Solving
di: Pan, Muyu, et al.
Pubblicazione: (2025)
di: Pan, Muyu, et al.
Pubblicazione: (2025)
Joint Verification and Refinement of Language Models for Safety-Constrained Planning
di: Yang, Yunhao, et al.
Pubblicazione: (2024)
di: Yang, Yunhao, et al.
Pubblicazione: (2024)
The Complexity of Aggregates over Extractions by Regular Expressions
di: Doleschal, Johannes, et al.
Pubblicazione: (2020)
di: Doleschal, Johannes, et al.
Pubblicazione: (2020)
Probabilistic Modeling of Spiking Neural Networks with Contract-Based Verification
di: Yao, Zhen, et al.
Pubblicazione: (2025)
di: Yao, Zhen, et al.
Pubblicazione: (2025)
BEAVER: An Efficient Deterministic LLM Verifier
di: Suresh, Tarun, et al.
Pubblicazione: (2025)
di: Suresh, Tarun, et al.
Pubblicazione: (2025)
Stochastic Directly-Follows Process Discovery Using Grammatical Inference
di: Alkhammash, Hanan, et al.
Pubblicazione: (2023)
di: Alkhammash, Hanan, et al.
Pubblicazione: (2023)
Inference of Deterministic Finite Automata via Q-Learning
di: Hosseinkhani, Elaheh, et al.
Pubblicazione: (2025)
di: Hosseinkhani, Elaheh, et al.
Pubblicazione: (2025)
Congruence-based Learning of Probabilistic Deterministic Finite Automata
di: Carrasco, Matías, et al.
Pubblicazione: (2024)
di: Carrasco, Matías, et al.
Pubblicazione: (2024)
Large Language Models and the Extended Church-Turing Thesis
di: Wiedermann, Jiří, et al.
Pubblicazione: (2024)
di: Wiedermann, Jiří, et al.
Pubblicazione: (2024)
RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance
di: Holzer, Nikolaus, et al.
Pubblicazione: (2025) -
Lost in Transmission: When and Why LLMs Fail to Reason Globally
di: Schnabel, Tobias, et al.
Pubblicazione: (2025) -
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
di: Wang, Ruida, et al.
Pubblicazione: (2024) -
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
di: Ospanov, Azim, et al.
Pubblicazione: (2025) -
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
di: Li, Zelong, et al.
Pubblicazione: (2024)