On Efficiently Representing Regular Languages as RNNs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Svete, Anej, Chan, Robin Shing Moon, Cotterell, Ryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformers Can Represent $n$-gram Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Unique Hard Attention: A Tale of Two Sides
von: Jerad, Selim, et al.
Veröffentlicht: (2025)
von: Jerad, Selim, et al.
Veröffentlicht: (2025)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
von: Svete, Anej, et al.
Veröffentlicht: (2026)
von: Svete, Anej, et al.
Veröffentlicht: (2026)
Context-Free Recognition with Transformers
von: Jerad, Selim, et al.
Veröffentlicht: (2026)
von: Jerad, Selim, et al.
Veröffentlicht: (2026)
On the Representational Capacity of Recurrent Neural Language Models
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
Gumbel Counterfactual Generation From Language Models
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2024)
Training Neural Networks as Recognizers of Formal Languages
von: Butoi, Alexandra, et al.
Veröffentlicht: (2024)
von: Butoi, Alexandra, et al.
Veröffentlicht: (2024)
On the Reasoning Abilities of Masked Diffusion Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2025)
von: Svete, Anej, et al.
Veröffentlicht: (2025)
Why Are Linear RNNs More Parallelizable?
von: Merrill, William, et al.
Veröffentlicht: (2026)
von: Merrill, William, et al.
Veröffentlicht: (2026)
What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages
von: Borenstein, Nadav, et al.
Veröffentlicht: (2024)
von: Borenstein, Nadav, et al.
Veröffentlicht: (2024)
On Affine Homotopy between Language Encoders
von: Chan, Robin SM, et al.
Veröffentlicht: (2024)
von: Chan, Robin SM, et al.
Veröffentlicht: (2024)
On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning
von: Nowak, Franz, et al.
Veröffentlicht: (2024)
von: Nowak, Franz, et al.
Veröffentlicht: (2024)
An $\mathbf{L^*}$ Algorithm for Deterministic Weighted Regular Languages
von: Pasti, Clemente, et al.
Veröffentlicht: (2024)
von: Pasti, Clemente, et al.
Veröffentlicht: (2024)
Lower Bounds on the Expressivity of Recurrent Neural Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Formal Aspects of Language Modeling
von: Cotterell, Ryan, et al.
Veröffentlicht: (2023)
von: Cotterell, Ryan, et al.
Veröffentlicht: (2023)
Can Transformers Learn $n$-gram Language Models?
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
A Geometric Notion of Causal Probing
von: Guerner, Clément, et al.
Veröffentlicht: (2023)
von: Guerner, Clément, et al.
Veröffentlicht: (2023)
Ensembling Language Models with Sequential Monte Carlo
von: Chan, Robin Shing Moon, et al.
Veröffentlicht: (2026)
von: Chan, Robin Shing Moon, et al.
Veröffentlicht: (2026)
Inference Scaling vs Reasoning: An Empirical Analysis of Compute-Optimal LLM Problem-Solving
von: AbdElhameed, Marwan, et al.
Veröffentlicht: (2024)
von: AbdElhameed, Marwan, et al.
Veröffentlicht: (2024)
Fundamental Limitations on Subquadratic Alternatives to Transformers
von: Alman, Josh, et al.
Veröffentlicht: (2024)
von: Alman, Josh, et al.
Veröffentlicht: (2024)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
von: Brösamle, Moritz, et al.
Veröffentlicht: (2026)
von: Brösamle, Moritz, et al.
Veröffentlicht: (2026)
Perfect diffusion is $\mathsf{TC}^0$ -- Bad diffusion is Turing-complete
von: Liu, Yuxi
Veröffentlicht: (2025)
von: Liu, Yuxi
Veröffentlicht: (2025)
The Role of $n$-gram Smoothing in the Age of Neural Networks
von: Malagutti, Luca, et al.
Veröffentlicht: (2024)
von: Malagutti, Luca, et al.
Veröffentlicht: (2024)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
von: Fan, Lizhou, et al.
Veröffentlicht: (2023)
On the Hardness of Learning Regular Expressions
von: Attias, Idan, et al.
Veröffentlicht: (2025)
von: Attias, Idan, et al.
Veröffentlicht: (2025)
Information Locality as an Inductive Bias for Neural Language Models
von: Someya, Taiga, et al.
Veröffentlicht: (2025)
von: Someya, Taiga, et al.
Veröffentlicht: (2025)
The Fine-Grained Complexity of Gradient Computation for Training Large Language Models
von: Alman, Josh, et al.
Veröffentlicht: (2024)
von: Alman, Josh, et al.
Veröffentlicht: (2024)
A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors
von: Tan, Naaman, et al.
Veröffentlicht: (2024)
von: Tan, Naaman, et al.
Veröffentlicht: (2024)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
RoPE Attention Can Be Trained in Almost Linear Time
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
von: Cao, Yang, et al.
Veröffentlicht: (2024)
von: Cao, Yang, et al.
Veröffentlicht: (2024)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
On Fine-Grained I/O Complexity of Attention Backward Passes
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
von: Daliri, Majid, et al.
Veröffentlicht: (2024)
von: Daliri, Majid, et al.
Veröffentlicht: (2024)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
von: Zhang, Yuyang, et al.
Veröffentlicht: (2026)
von: Zhang, Yuyang, et al.
Veröffentlicht: (2026)
Demystifying the unreasonable effectiveness of online alignment methods
von: Kang, Enoch Hyunwook
Veröffentlicht: (2026)
von: Kang, Enoch Hyunwook
Veröffentlicht: (2026)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
von: Huang, Zekai, et al.
Veröffentlicht: (2025)
von: Huang, Zekai, et al.
Veröffentlicht: (2025)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transformers Can Represent $n$-gram Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2024) -
Unique Hard Attention: A Tale of Two Sides
von: Jerad, Selim, et al.
Veröffentlicht: (2025) -
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
von: Svete, Anej, et al.
Veröffentlicht: (2026) -
Context-Free Recognition with Transformers
von: Jerad, Selim, et al.
Veröffentlicht: (2026) -
On the Representational Capacity of Recurrent Neural Language Models
von: Nowak, Franz, et al.
Veröffentlicht: (2023)