Probability Distributions Computed by Autoregressive Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Andy, Svete, Anej, Li, Jiaoda, Lin, Anthony Widjaja, Rawski, Jonathan, Cotterell, Ryan, Chiang, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unique Hard Attention: A Tale of Two Sides
by: Jerad, Selim, et al.
Published: (2025)
by: Jerad, Selim, et al.
Published: (2025)
Transformers Can Represent $n$-gram Language Models
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Characterizing the Expressivity of Local Attention in Transformers
by: Li, Jiaoda, et al.
Published: (2026)
by: Li, Jiaoda, et al.
Published: (2026)
Characterizing the Expressivity of Fixed-Precision Transformer Language Models
by: Li, Jiaoda, et al.
Published: (2025)
by: Li, Jiaoda, et al.
Published: (2025)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
by: Svete, Anej, et al.
Published: (2026)
by: Svete, Anej, et al.
Published: (2026)
Context-Free Recognition with Transformers
by: Jerad, Selim, et al.
Published: (2026)
by: Jerad, Selim, et al.
Published: (2026)
Can Transformers Learn $n$-gram Language Models?
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
On the Representational Capacity of Recurrent Neural Language Models
by: Nowak, Franz, et al.
Published: (2023)
by: Nowak, Franz, et al.
Published: (2023)
On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning
by: Nowak, Franz, et al.
Published: (2024)
by: Nowak, Franz, et al.
Published: (2024)
On Efficiently Representing Regular Languages as RNNs
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Lower Bounds on the Expressivity of Recurrent Neural Language Models
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Gumbel Counterfactual Generation From Language Models
by: Ravfogel, Shauli, et al.
Published: (2024)
by: Ravfogel, Shauli, et al.
Published: (2024)
Formal Aspects of Language Modeling
by: Cotterell, Ryan, et al.
Published: (2023)
by: Cotterell, Ryan, et al.
Published: (2023)
A Geometric Notion of Causal Probing
by: Guerner, Clément, et al.
Published: (2023)
by: Guerner, Clément, et al.
Published: (2023)
A Probability--Quality Trade-off in Aligned Language Models and its Relation to Sampling Adaptors
by: Tan, Naaman, et al.
Published: (2024)
by: Tan, Naaman, et al.
Published: (2024)
A Transformer with Stack Attention
by: Li, Jiaoda, et al.
Published: (2024)
by: Li, Jiaoda, et al.
Published: (2024)
The Role of $n$-gram Smoothing in the Age of Neural Networks
by: Malagutti, Luca, et al.
Published: (2024)
by: Malagutti, Luca, et al.
Published: (2024)
An $\mathbf{L^*}$ Algorithm for Deterministic Weighted Regular Languages
by: Pasti, Clemente, et al.
Published: (2024)
by: Pasti, Clemente, et al.
Published: (2024)
On the Reasoning Abilities of Masked Diffusion Language Models
by: Svete, Anej, et al.
Published: (2025)
by: Svete, Anej, et al.
Published: (2025)
Training Neural Networks as Recognizers of Formal Languages
by: Butoi, Alexandra, et al.
Published: (2024)
by: Butoi, Alexandra, et al.
Published: (2024)
What Do Language Models Learn in Context? The Structured Task Hypothesis
by: Li, Jiaoda, et al.
Published: (2024)
by: Li, Jiaoda, et al.
Published: (2024)
Information Locality as an Inductive Bias for Neural Language Models
by: Someya, Taiga, et al.
Published: (2025)
by: Someya, Taiga, et al.
Published: (2025)
What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages
by: Borenstein, Nadav, et al.
Published: (2024)
by: Borenstein, Nadav, et al.
Published: (2024)
Counting Like Transformers: Compiling Temporal Counting Logic Into Softmax Transformers
by: Yang, Andy, et al.
Published: (2024)
by: Yang, Andy, et al.
Published: (2024)
Transformers are Inherently Succinct
by: Bergsträßer, Pascal, et al.
Published: (2025)
by: Bergsträßer, Pascal, et al.
Published: (2025)
Transformers as Transducers
by: Strobl, Lena, et al.
Published: (2024)
by: Strobl, Lena, et al.
Published: (2024)
A Fast Algorithm for Computing Prefix Probabilities
by: Nowak, Franz, et al.
Published: (2023)
by: Nowak, Franz, et al.
Published: (2023)
Knee-Deep in C-RASP: A Transformer Depth Hierarchy
by: Yang, Andy, et al.
Published: (2025)
by: Yang, Andy, et al.
Published: (2025)
On Affine Homotopy between Language Encoders
by: Chan, Robin SM, et al.
Published: (2024)
by: Chan, Robin SM, et al.
Published: (2024)
The Counting Power of Transformers
by: Sälzer, Marco, et al.
Published: (2025)
by: Sälzer, Marco, et al.
Published: (2025)
Length Generalization Bounds for Transformers
by: Yang, Andy, et al.
Published: (2026)
by: Yang, Andy, et al.
Published: (2026)
Algorithms for Weighted Pushdown Automata
by: Butoi, Alexandra, et al.
Published: (2022)
by: Butoi, Alexandra, et al.
Published: (2022)
Softmax Transformers are Turing-Complete
by: Jiang, Hongjian, et al.
Published: (2025)
by: Jiang, Hongjian, et al.
Published: (2025)
Synthesis and Verification of Transformer Programs (Technical Report)
by: Jiang, Hongjian, et al.
Published: (2026)
by: Jiang, Hongjian, et al.
Published: (2026)
A Distributional Perspective on Word Learning in Neural Language Models
by: Ficarra, Filippo, et al.
Published: (2025)
by: Ficarra, Filippo, et al.
Published: (2025)
Efficiently Computing Susceptibility to Context in Language Models
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
Exact Hard Monotonic Attention for Character-Level Transduction
by: Wu, Shijie, et al.
Published: (2019)
by: Wu, Shijie, et al.
Published: (2019)
Low-Resource Named Entity Recognition with Cross-Lingual, Character-Level Neural Conditional Random Fields
by: Cotterell, Ryan, et al.
Published: (2024)
by: Cotterell, Ryan, et al.
Published: (2024)
Cross-lingual, Character-Level Neural Morphological Tagging
by: Cotterell, Ryan, et al.
Published: (2017)
by: Cotterell, Ryan, et al.
Published: (2017)
Masked Hard-Attention Transformers Recognize Exactly the Star-Free Languages
by: Yang, Andy, et al.
Published: (2023)
by: Yang, Andy, et al.
Published: (2023)
Similar Items
-
Unique Hard Attention: A Tale of Two Sides
by: Jerad, Selim, et al.
Published: (2025) -
Transformers Can Represent $n$-gram Language Models
by: Svete, Anej, et al.
Published: (2024) -
Characterizing the Expressivity of Local Attention in Transformers
by: Li, Jiaoda, et al.
Published: (2026) -
Characterizing the Expressivity of Fixed-Precision Transformer Language Models
by: Li, Jiaoda, et al.
Published: (2025) -
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
by: Svete, Anej, et al.
Published: (2026)