Probability Distributions Computed by Autoregressive Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916041889677312 |
|---|---|
| author | Yang, Andy Svete, Anej Li, Jiaoda Lin, Anthony Widjaja Rawski, Jonathan Cotterell, Ryan Chiang, David |
| author_facet | Yang, Andy Svete, Anej Li, Jiaoda Lin, Anthony Widjaja Rawski, Jonathan Cotterell, Ryan Chiang, David |
| contents | Most expressivity results for transformers treat them as language recognizers -- devices that accept or reject strings -- rather than as they are used in practice: as language models that generate strings autoregressively and probabilistically. We characterize the probability distributions that transformer language models can express. We show that making transformer language recognizers autoregressive can sometimes increase their expressivity, and that making them probabilistic can break equivalences that hold in the non-probabilistic case. Our overall contribution is to tease apart what functions transformers are capable of expressing in their most common use case as language models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_27118 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Probability Distributions Computed by Autoregressive Transformers Yang, Andy Svete, Anej Li, Jiaoda Lin, Anthony Widjaja Rawski, Jonathan Cotterell, Ryan Chiang, David Computation and Language Most expressivity results for transformers treat them as language recognizers -- devices that accept or reject strings -- rather than as they are used in practice: as language models that generate strings autoregressively and probabilistically. We characterize the probability distributions that transformer language models can express. We show that making transformer language recognizers autoregressive can sometimes increase their expressivity, and that making them probabilistic can break equivalences that hold in the non-probabilistic case. Our overall contribution is to tease apart what functions transformers are capable of expressing in their most common use case as language models. |
| title | Probability Distributions Computed by Autoregressive Transformers |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.27118 |