Learning State-Tracking from Code Using Linear RNNs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Siems, Julien, Grazzi, Riccardo, Kalinin, Kirill, Ballani, Hitesh, Rahmani, Babak
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918463778324480
author Siems, Julien
Grazzi, Riccardo
Kalinin, Kirill
Ballani, Hitesh
Rahmani, Babak
author_facet Siems, Julien
Grazzi, Riccardo
Kalinin, Kirill
Ballani, Hitesh
Rahmani, Babak
contents Over the last years, state-tracking tasks, particularly permutation composition, have become a testbed to understand the limits of sequence models architectures like Transformers and RNNs (linear and non-linear). However, these are often sequence-to-sequence tasks: learning to map actions (permutations) to states, which is incompatible with the next-token prediction setting commonly used to train language models. We address this gap by converting permutation composition into code via REPL traces that interleave state-reveals through prints and variable transformations. We show that linear RNNs capable of state-tracking excel also in this setting, while Transformers still fail. Motivated by this representation, we investigate why tracking states in code is generally difficult: actions are not always fully observable. We frame this as tracking the state of a probabilistic finite-state automaton with deterministic state reveals and show that linear RNNs can be worse than non-linear RNNs at tracking states in this setup.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14814
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning State-Tracking from Code Using Linear RNNs
Siems, Julien
Grazzi, Riccardo
Kalinin, Kirill
Ballani, Hitesh
Rahmani, Babak
Machine Learning
Computation and Language
Over the last years, state-tracking tasks, particularly permutation composition, have become a testbed to understand the limits of sequence models architectures like Transformers and RNNs (linear and non-linear). However, these are often sequence-to-sequence tasks: learning to map actions (permutations) to states, which is incompatible with the next-token prediction setting commonly used to train language models. We address this gap by converting permutation composition into code via REPL traces that interleave state-reveals through prints and variable transformations. We show that linear RNNs capable of state-tracking excel also in this setting, while Transformers still fail. Motivated by this representation, we investigate why tracking states in code is generally difficult: actions are not always fully observable. We frame this as tracking the state of a probabilistic finite-state automaton with deterministic state reveals and show that linear RNNs can be worse than non-linear RNNs at tracking states in this setup.
title Learning State-Tracking from Code Using Linear RNNs
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2602.14814