Characterizing the Expressivity of Fixed-Precision Transformer Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiaoda, Cotterell, Ryan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Characterizing the Expressivity of Local Attention in Transformers
by: Li, Jiaoda, et al.
Published: (2026)
by: Li, Jiaoda, et al.
Published: (2026)
What Do Language Models Learn in Context? The Structured Task Hypothesis
by: Li, Jiaoda, et al.
Published: (2024)
by: Li, Jiaoda, et al.
Published: (2024)
A Transformer with Stack Attention
by: Li, Jiaoda, et al.
Published: (2024)
by: Li, Jiaoda, et al.
Published: (2024)
Unique Hard Attention: A Tale of Two Sides
by: Jerad, Selim, et al.
Published: (2025)
by: Jerad, Selim, et al.
Published: (2025)
An Algebraic View of the Expressivity of Recurrent Language Models
by: Nowak, Franz, et al.
Published: (2026)
by: Nowak, Franz, et al.
Published: (2026)
Probability Distributions Computed by Autoregressive Transformers
by: Yang, Andy, et al.
Published: (2025)
by: Yang, Andy, et al.
Published: (2025)
Lower Bounds on the Expressivity of Recurrent Neural Language Models
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Transformers Can Represent $n$-gram Language Models
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
by: Svete, Anej, et al.
Published: (2026)
by: Svete, Anej, et al.
Published: (2026)
Can Transformers Learn $n$-gram Language Models?
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Formal Aspects of Language Modeling
by: Cotterell, Ryan, et al.
Published: (2023)
by: Cotterell, Ryan, et al.
Published: (2023)
Investigating Critical Period Effects in Language Acquisition through Neural Language Models
by: Constantinescu, Ionut, et al.
Published: (2024)
by: Constantinescu, Ionut, et al.
Published: (2024)
A Distributional Perspective on Word Learning in Neural Language Models
by: Ficarra, Filippo, et al.
Published: (2025)
by: Ficarra, Filippo, et al.
Published: (2025)
Efficiently Computing Susceptibility to Context in Language Models
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
On the Representational Capacity of Recurrent Neural Language Models
by: Nowak, Franz, et al.
Published: (2023)
by: Nowak, Franz, et al.
Published: (2023)
Towards Explainability in Legal Outcome Prediction Models
by: Valvoda, Josef, et al.
Published: (2024)
by: Valvoda, Josef, et al.
Published: (2024)
On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning
by: Nowak, Franz, et al.
Published: (2024)
by: Nowak, Franz, et al.
Published: (2024)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
by: Amini, Afra, et al.
Published: (2025)
by: Amini, Afra, et al.
Published: (2025)
Exact Hard Monotonic Attention for Character-Level Transduction
by: Wu, Shijie, et al.
Published: (2019)
by: Wu, Shijie, et al.
Published: (2019)
Low-Resource Named Entity Recognition with Cross-Lingual, Character-Level Neural Conditional Random Fields
by: Cotterell, Ryan, et al.
Published: (2024)
by: Cotterell, Ryan, et al.
Published: (2024)
Cross-lingual, Character-Level Neural Morphological Tagging
by: Cotterell, Ryan, et al.
Published: (2017)
by: Cotterell, Ryan, et al.
Published: (2017)
A Simple Joint Model for Improved Contextual Neural Lemmatization
by: Malaviya, Chaitanya, et al.
Published: (2019)
by: Malaviya, Chaitanya, et al.
Published: (2019)
Transducing Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2026)
by: Snæbjarnarson, Vésteinn, et al.
Published: (2026)
Context-Free Recognition with Transformers
by: Jerad, Selim, et al.
Published: (2026)
by: Jerad, Selim, et al.
Published: (2026)
Can Language Models Learn Typologically Implausible Languages?
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
by: Vargas, Francisco, et al.
Published: (2020)
by: Vargas, Francisco, et al.
Published: (2020)
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
by: DuSell, Brian, et al.
Published: (2025)
by: DuSell, Brian, et al.
Published: (2025)
Syntactic Control of Language Models by Posterior Inference
by: Xefteri, Vicky, et al.
Published: (2025)
by: Xefteri, Vicky, et al.
Published: (2025)
Gumbel Counterfactual Generation From Language Models
by: Ravfogel, Shauli, et al.
Published: (2024)
by: Ravfogel, Shauli, et al.
Published: (2024)
Towards Developmentally Plausible Rewards: Communicative Success as a Learning Signal for Interactive Language Models
by: Stöpler, Lennart, et al.
Published: (2025)
by: Stöpler, Lennart, et al.
Published: (2025)
Generalized Measures of Anticipation and Responsivity in Online Language Processing
by: Giulianelli, Mario, et al.
Published: (2024)
by: Giulianelli, Mario, et al.
Published: (2024)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
by: Brösamle, Moritz, et al.
Published: (2026)
by: Brösamle, Moritz, et al.
Published: (2026)
On Efficiently Representing Regular Languages as RNNs
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Transformers are Inherently Succinct
by: Bergsträßer, Pascal, et al.
Published: (2025)
by: Bergsträßer, Pascal, et al.
Published: (2025)
Labeled Morphological Segmentation with Semi-Markov Models
by: Cotterell, Ryan, et al.
Published: (2024)
by: Cotterell, Ryan, et al.
Published: (2024)
Context versus Prior Knowledge in Language Models
by: Du, Kevin, et al.
Published: (2024)
by: Du, Kevin, et al.
Published: (2024)
Structured Voronoi Sampling
by: Amini, Afra, et al.
Published: (2023)
by: Amini, Afra, et al.
Published: (2023)
Hard Non-Monotonic Attention for Character-Level Transduction
by: Wu, Shijie, et al.
Published: (2018)
by: Wu, Shijie, et al.
Published: (2018)
Activation Scaling for Steering and Interpreting Language Models
by: Stoehr, Niklas, et al.
Published: (2024)
by: Stoehr, Niklas, et al.
Published: (2024)
Post-Training Language Models for Crosslingual Consistency
by: Liu, Tianyu, et al.
Published: (2026)
by: Liu, Tianyu, et al.
Published: (2026)
Similar Items
-
Characterizing the Expressivity of Local Attention in Transformers
by: Li, Jiaoda, et al.
Published: (2026) -
What Do Language Models Learn in Context? The Structured Task Hypothesis
by: Li, Jiaoda, et al.
Published: (2024) -
A Transformer with Stack Attention
by: Li, Jiaoda, et al.
Published: (2024) -
Unique Hard Attention: A Tale of Two Sides
by: Jerad, Selim, et al.
Published: (2025) -
An Algebraic View of the Expressivity of Recurrent Language Models
by: Nowak, Franz, et al.
Published: (2026)