Simulating Hard Attention Using Soft Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Andy, Strobl, Lena, Chiang, David, Angluin, Dana |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
by: Yang, Andy, et al.
Published: (2023)
by: Yang, Andy, et al.
Published: (2023)
Transformers as Transducers
by: Strobl, Lena, et al.
Published: (2024)
by: Strobl, Lena, et al.
Published: (2024)
What Formal Languages Can Transformers Express? A Survey
by: Strobl, Lena, et al.
Published: (2023)
by: Strobl, Lena, et al.
Published: (2023)
Constructing Concise Characteristic Samples for Acceptors of Omega Regular Languages
by: Angluin, Dana, et al.
Published: (2022)
by: Angluin, Dana, et al.
Published: (2022)
Unique Hard Attention: A Tale of Two Sides
by: Jerad, Selim, et al.
Published: (2025)
by: Jerad, Selim, et al.
Published: (2025)
Counting Like Transformers: Compiling Temporal Counting Logic Into Softmax Transformers
by: Yang, Andy, et al.
Published: (2024)
by: Yang, Andy, et al.
Published: (2024)
Knee-Deep in C-RASP: A Transformer Depth Hierarchy
by: Yang, Andy, et al.
Published: (2025)
by: Yang, Andy, et al.
Published: (2025)
Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
by: Awadhiya, Kanishk
Published: (2026)
by: Awadhiya, Kanishk
Published: (2026)
Transformers in Uniform TC$^0$
by: Chiang, David
Published: (2024)
by: Chiang, David
Published: (2024)
Length Generalization Bounds for Transformers
by: Yang, Andy, et al.
Published: (2026)
by: Yang, Andy, et al.
Published: (2026)
The Power of Hard Attention Transformers on Data Sequences: A Formal Language Theoretic Perspective
by: Bergsträßer, Pascal, et al.
Published: (2024)
by: Bergsträßer, Pascal, et al.
Published: (2024)
The Expressive Capacity of State Space Models: A Formal Language Perspective
by: Sarrof, Yash, et al.
Published: (2024)
by: Sarrof, Yash, et al.
Published: (2024)
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
by: Grazzi, Riccardo, et al.
Published: (2024)
by: Grazzi, Riccardo, et al.
Published: (2024)
Constructing a BPE Tokenization DFA
by: Berglund, Martin, et al.
Published: (2024)
by: Berglund, Martin, et al.
Published: (2024)
Language Models over Canonical Byte-Pair Encodings
by: Vieira, Tim, et al.
Published: (2025)
by: Vieira, Tim, et al.
Published: (2025)
An Algebraic View of the Expressivity of Recurrent Language Models
by: Nowak, Franz, et al.
Published: (2026)
by: Nowak, Franz, et al.
Published: (2026)
From Formal Language Theory to Statistical Learning: Finite Observability of Subregular Languages
by: Hayashi, Katsuhiko, et al.
Published: (2025)
by: Hayashi, Katsuhiko, et al.
Published: (2025)
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
by: Siems, Julien, et al.
Published: (2025)
by: Siems, Julien, et al.
Published: (2025)
Sampling from Your Language Model One Byte at a Time
by: Hayase, Jonathan, et al.
Published: (2025)
by: Hayase, Jonathan, et al.
Published: (2025)
Unraveling Syntax: How Language Models Learn Context-Free Grammars
by: Schulz, Laura Ying, et al.
Published: (2025)
by: Schulz, Laura Ying, et al.
Published: (2025)
Comparison of different Unique hard attention transformer models by the formal languages they can recognize
by: Ryvkin, Leonid
Published: (2025)
by: Ryvkin, Leonid
Published: (2025)
Correct and Optimal: the Regular Expression Inference Challenge
by: Valizadeh, Mojtaba, et al.
Published: (2023)
by: Valizadeh, Mojtaba, et al.
Published: (2023)
The Counting Power of Transformers
by: Sälzer, Marco, et al.
Published: (2025)
by: Sälzer, Marco, et al.
Published: (2025)
MLRegTest: A Benchmark for the Machine Learning of Regular Languages
by: van der Poel, Sam, et al.
Published: (2023)
by: van der Poel, Sam, et al.
Published: (2023)
CoT-TL: Low-Resource Temporal Knowledge Representation of Planning Instructions Using Chain-of-Thought Reasoning
by: Manas, Kumar, et al.
Published: (2024)
by: Manas, Kumar, et al.
Published: (2024)
Atomic Gliders and CA as Language Generators (Extended Version)
by: Fisman, Dana, et al.
Published: (2025)
by: Fisman, Dana, et al.
Published: (2025)
Learning Weighted Finite Automata over the Max-Plus Semiring and its Termination
by: Okudono, Takamasa, et al.
Published: (2024)
by: Okudono, Takamasa, et al.
Published: (2024)
PDFA Distillation via String Probability Queries
by: Baumgartner, Robert, et al.
Published: (2024)
by: Baumgartner, Robert, et al.
Published: (2024)
Certifying Robustness of Graph Convolutional Networks for Node Perturbation with Polyhedra Abstract Interpretation
by: Chen, Boqi, et al.
Published: (2024)
by: Chen, Boqi, et al.
Published: (2024)
Extending AALpy with Passive Learning: A Generalized State-Merging Approach
by: von Berg, Benjamin, et al.
Published: (2025)
by: von Berg, Benjamin, et al.
Published: (2025)
Finite Sentence-Interface Control for Learning Bounded-Fan-Out Linear MCFGs under Fixed Monoid Typing
by: Kuriyama, Takayuki
Published: (2026)
by: Kuriyama, Takayuki
Published: (2026)
Deconstructing Subset Construction -- Reducing While Determinizing
by: Nicol, John, et al.
Published: (2025)
by: Nicol, John, et al.
Published: (2025)
Learning Reward Machines from Partially Observed Policies
by: Shehab, Mohamad Louai, et al.
Published: (2025)
by: Shehab, Mohamad Louai, et al.
Published: (2025)
Stochastic Alignments: Matching an Observed Trace to Stochastic Process Models
by: Li, Tian, et al.
Published: (2025)
by: Li, Tian, et al.
Published: (2025)
A Constructive Framework for Nondeterministic Automata via Time-Shared, Depth-Unrolled Feedforward Networks
by: Dhayalkar, Sahil Rajesh
Published: (2025)
by: Dhayalkar, Sahil Rajesh
Published: (2025)
Active Learning of Symbolic Automata Over Rational Numbers
by: Hagedorn, Sebastian, et al.
Published: (2025)
by: Hagedorn, Sebastian, et al.
Published: (2025)
Learning Deterministic Finite-State Machines from the Prefixes of a Single String is NP-Complete
by: Dumitru, Radu Cosmin, et al.
Published: (2026)
by: Dumitru, Radu Cosmin, et al.
Published: (2026)
SMT-Based Active Learning of Weighted Automata
by: Ferreira, Tiago, et al.
Published: (2026)
by: Ferreira, Tiago, et al.
Published: (2026)
Continuous Diffusion Models Can Obey Formal Syntax
by: Kim, Jinwoo, et al.
Published: (2026)
by: Kim, Jinwoo, et al.
Published: (2026)
Unsupervised Hierarchical Skill Discovery
by: Harvey, Damion, et al.
Published: (2026)
by: Harvey, Damion, et al.
Published: (2026)
Similar Items
-
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
by: Yang, Andy, et al.
Published: (2023) -
Transformers as Transducers
by: Strobl, Lena, et al.
Published: (2024) -
What Formal Languages Can Transformers Express? A Survey
by: Strobl, Lena, et al.
Published: (2023) -
Constructing Concise Characteristic Samples for Acceptors of Omega Regular Languages
by: Angluin, Dana, et al.
Published: (2022) -
Unique Hard Attention: A Tale of Two Sides
by: Jerad, Selim, et al.
Published: (2025)