Emergent Stack Representations in Modeling Counter Languages Using Transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tiwari, Utkarsh, Gupta, Aviral, Hahn, Michael
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909473483784192
author Tiwari, Utkarsh
Gupta, Aviral
Hahn, Michael
author_facet Tiwari, Utkarsh
Gupta, Aviral
Hahn, Michael
contents Transformer architectures are the backbone of most modern language models, but understanding the inner workings of these models still largely remains an open problem. One way that research in the past has tackled this problem is by isolating the learning capabilities of these architectures by training them over well-understood classes of formal languages. We extend this literature by analyzing models trained over counter languages, which can be modeled using counter variables. We train transformer models on 4 counter languages, and equivalently formulate these languages using stacks, whose depths can be understood as the counter values. We then probe their internal representations for stack depths at each input token to show that these models when trained as next token predictors learn stack-like representations. This brings us closer to understanding the algorithmic details of how transformers learn languages and helps in circuit discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Emergent Stack Representations in Modeling Counter Languages Using Transformers
Tiwari, Utkarsh
Gupta, Aviral
Hahn, Michael
Computation and Language
Machine Learning
Transformer architectures are the backbone of most modern language models, but understanding the inner workings of these models still largely remains an open problem. One way that research in the past has tackled this problem is by isolating the learning capabilities of these architectures by training them over well-understood classes of formal languages. We extend this literature by analyzing models trained over counter languages, which can be modeled using counter variables. We train transformer models on 4 counter languages, and equivalently formulate these languages using stacks, whose depths can be understood as the counter values. We then probe their internal representations for stack depths at each input token to show that these models when trained as next token predictors learn stack-like representations. This brings us closer to understanding the algorithmic details of how transformers learn languages and helps in circuit discovery.
title Emergent Stack Representations in Modeling Counter Languages Using Transformers
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.01432