Probability Distributions Computed by Autoregressive Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Andy, Svete, Anej, Li, Jiaoda, Lin, Anthony Widjaja, Rawski, Jonathan, Cotterell, Ryan, Chiang, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916041889677312
author Yang, Andy
Svete, Anej
Li, Jiaoda
Lin, Anthony Widjaja
Rawski, Jonathan
Cotterell, Ryan
Chiang, David
author_facet Yang, Andy
Svete, Anej
Li, Jiaoda
Lin, Anthony Widjaja
Rawski, Jonathan
Cotterell, Ryan
Chiang, David
contents Most expressivity results for transformers treat them as language recognizers -- devices that accept or reject strings -- rather than as they are used in practice: as language models that generate strings autoregressively and probabilistically. We characterize the probability distributions that transformer language models can express. We show that making transformer language recognizers autoregressive can sometimes increase their expressivity, and that making them probabilistic can break equivalences that hold in the non-probabilistic case. Our overall contribution is to tease apart what functions transformers are capable of expressing in their most common use case as language models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27118
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probability Distributions Computed by Autoregressive Transformers
Yang, Andy
Svete, Anej
Li, Jiaoda
Lin, Anthony Widjaja
Rawski, Jonathan
Cotterell, Ryan
Chiang, David
Computation and Language
Most expressivity results for transformers treat them as language recognizers -- devices that accept or reject strings -- rather than as they are used in practice: as language models that generate strings autoregressively and probabilistically. We characterize the probability distributions that transformer language models can express. We show that making transformer language recognizers autoregressive can sometimes increase their expressivity, and that making them probabilistic can break equivalences that hold in the non-probabilistic case. Our overall contribution is to tease apart what functions transformers are capable of expressing in their most common use case as language models.
title Probability Distributions Computed by Autoregressive Transformers
topic Computation and Language
url https://arxiv.org/abs/2510.27118