Characterizing the Expressivity of Local Attention in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jiaoda, Cotterell, Ryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Characterizing the Expressivity of Fixed-Precision Transformer Language Models
von: Li, Jiaoda, et al.
Veröffentlicht: (2025)
von: Li, Jiaoda, et al.
Veröffentlicht: (2025)
A Transformer with Stack Attention
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
Unique Hard Attention: A Tale of Two Sides
von: Jerad, Selim, et al.
Veröffentlicht: (2025)
von: Jerad, Selim, et al.
Veröffentlicht: (2025)
What Do Language Models Learn in Context? The Structured Task Hypothesis
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
Probability Distributions Computed by Autoregressive Transformers
von: Yang, Andy, et al.
Veröffentlicht: (2025)
von: Yang, Andy, et al.
Veröffentlicht: (2025)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
von: Svete, Anej, et al.
Veröffentlicht: (2026)
von: Svete, Anej, et al.
Veröffentlicht: (2026)
An Algebraic View of the Expressivity of Recurrent Language Models
von: Nowak, Franz, et al.
Veröffentlicht: (2026)
von: Nowak, Franz, et al.
Veröffentlicht: (2026)
Exact Hard Monotonic Attention for Character-Level Transduction
von: Wu, Shijie, et al.
Veröffentlicht: (2019)
von: Wu, Shijie, et al.
Veröffentlicht: (2019)
Lower Bounds on the Expressivity of Recurrent Neural Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Transformers Can Represent $n$-gram Language Models
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Hard Non-Monotonic Attention for Character-Level Transduction
von: Wu, Shijie, et al.
Veröffentlicht: (2018)
von: Wu, Shijie, et al.
Veröffentlicht: (2018)
Low-Resource Named Entity Recognition with Cross-Lingual, Character-Level Neural Conditional Random Fields
von: Cotterell, Ryan, et al.
Veröffentlicht: (2024)
von: Cotterell, Ryan, et al.
Veröffentlicht: (2024)
Cross-lingual, Character-Level Neural Morphological Tagging
von: Cotterell, Ryan, et al.
Veröffentlicht: (2017)
von: Cotterell, Ryan, et al.
Veröffentlicht: (2017)
Locally Typical Sampling
von: Meister, Clara, et al.
Veröffentlicht: (2022)
von: Meister, Clara, et al.
Veröffentlicht: (2022)
Context-Free Recognition with Transformers
von: Jerad, Selim, et al.
Veröffentlicht: (2026)
von: Jerad, Selim, et al.
Veröffentlicht: (2026)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
von: DuSell, Brian, et al.
Veröffentlicht: (2025)
von: DuSell, Brian, et al.
Veröffentlicht: (2025)
Towards Explainability in Legal Outcome Prediction Models
von: Valvoda, Josef, et al.
Veröffentlicht: (2024)
von: Valvoda, Josef, et al.
Veröffentlicht: (2024)
Can Transformers Learn $n$-gram Language Models?
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Transformers are Inherently Succinct
von: Bergsträßer, Pascal, et al.
Veröffentlicht: (2025)
von: Bergsträßer, Pascal, et al.
Veröffentlicht: (2025)
Structured Voronoi Sampling
von: Amini, Afra, et al.
Veröffentlicht: (2023)
von: Amini, Afra, et al.
Veröffentlicht: (2023)
A Simple Joint Model for Improved Contextual Neural Lemmatization
von: Malaviya, Chaitanya, et al.
Veröffentlicht: (2019)
von: Malaviya, Chaitanya, et al.
Veröffentlicht: (2019)
Log-linear Guardedness and its Implications
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
A Distributional Perspective on Word Learning in Neural Language Models
von: Ficarra, Filippo, et al.
Veröffentlicht: (2025)
von: Ficarra, Filippo, et al.
Veröffentlicht: (2025)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2024)
Information Locality as an Inductive Bias for Neural Language Models
von: Someya, Taiga, et al.
Veröffentlicht: (2025)
von: Someya, Taiga, et al.
Veröffentlicht: (2025)
On the Proper Treatment of Units in Surprisal Theory
von: Kiegeland, Samuel, et al.
Veröffentlicht: (2026)
von: Kiegeland, Samuel, et al.
Veröffentlicht: (2026)
Efficiently Computing Susceptibility to Context in Language Models
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
Joint Lemmatization and Morphological Tagging with LEMMING
von: Muller, Thomas, et al.
Veröffentlicht: (2024)
von: Muller, Thomas, et al.
Veröffentlicht: (2024)
Investigating Critical Period Effects in Language Acquisition through Neural Language Models
von: Constantinescu, Ionut, et al.
Veröffentlicht: (2024)
von: Constantinescu, Ionut, et al.
Veröffentlicht: (2024)
Labeled Morphological Segmentation with Semi-Markov Models
von: Cotterell, Ryan, et al.
Veröffentlicht: (2024)
von: Cotterell, Ryan, et al.
Veröffentlicht: (2024)
On the Representational Capacity of Recurrent Neural Language Models
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
von: Nowak, Franz, et al.
Veröffentlicht: (2023)
On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning
von: Nowak, Franz, et al.
Veröffentlicht: (2024)
von: Nowak, Franz, et al.
Veröffentlicht: (2024)
Formal Aspects of Language Modeling
von: Cotterell, Ryan, et al.
Veröffentlicht: (2023)
von: Cotterell, Ryan, et al.
Veröffentlicht: (2023)
More Expressive Attention with Negative Weights
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
Kimi Linear: An Expressive, Efficient Attention Architecture
von: Kimi Team, et al.
Veröffentlicht: (2025)
von: Kimi Team, et al.
Veröffentlicht: (2025)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
von: Amini, Afra, et al.
Veröffentlicht: (2025)
von: Amini, Afra, et al.
Veröffentlicht: (2025)
Direct Preference Optimization with an Offset
von: Amini, Afra, et al.
Veröffentlicht: (2024)
von: Amini, Afra, et al.
Veröffentlicht: (2024)
Generalized Measures of Anticipation and Responsivity in Online Language Processing
von: Giulianelli, Mario, et al.
Veröffentlicht: (2024)
von: Giulianelli, Mario, et al.
Veröffentlicht: (2024)
Speakers Fill Lexical Semantic Gaps with Context
von: Pimentel, Tiago, et al.
Veröffentlicht: (2020)
von: Pimentel, Tiago, et al.
Veröffentlicht: (2020)
Ähnliche Einträge
-
Characterizing the Expressivity of Fixed-Precision Transformer Language Models
von: Li, Jiaoda, et al.
Veröffentlicht: (2025) -
A Transformer with Stack Attention
von: Li, Jiaoda, et al.
Veröffentlicht: (2024) -
Unique Hard Attention: A Tale of Two Sides
von: Jerad, Selim, et al.
Veröffentlicht: (2025) -
What Do Language Models Learn in Context? The Structured Task Hypothesis
von: Li, Jiaoda, et al.
Veröffentlicht: (2024) -
Probability Distributions Computed by Autoregressive Transformers
von: Yang, Andy, et al.
Veröffentlicht: (2025)