The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Edelman, Benjamin L., Edelman, Ezra, Goel, Surbhi, Malach, Eran, Tsilivis, Nikolaos
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910334779916288
author Edelman, Benjamin L.
Edelman, Ezra
Goel, Surbhi
Malach, Eran
Tsilivis, Nikolaos
author_facet Edelman, Benjamin L.
Edelman, Ezra
Goel, Surbhi
Malach, Eran
Tsilivis, Nikolaos
contents Large language models have the ability to generate text that mimics patterns in their inputs. We introduce a simple Markov Chain sequence modeling task in order to study how this in-context learning (ICL) capability emerges. In our setting, each example is sampled from a Markov chain drawn from a prior distribution over Markov chains. Transformers trained on this task form \emph{statistical induction heads} which compute accurate next-token probabilities given the bigram statistics of the context. During the course of training, models pass through multiple phases: after an initial stage in which predictions are uniform, they learn to sub-optimally predict using in-context single-token statistics (unigrams); then, there is a rapid phase transition to the correct in-context bigram solution. We conduct an empirical and theoretical investigation of this multi-phase process, showing how successful learning results from the interaction between the transformer's layers, and uncovering evidence that the presence of the simpler unigram solution may delay formation of the final bigram solution. We examine how learning is affected by varying the prior distribution over Markov chains, and consider the generalization of our in-context learning of Markov chains (ICL-MC) task to $n$-grams for $n > 2$.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11004
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
Edelman, Benjamin L.
Edelman, Ezra
Goel, Surbhi
Malach, Eran
Tsilivis, Nikolaos
Machine Learning
Large language models have the ability to generate text that mimics patterns in their inputs. We introduce a simple Markov Chain sequence modeling task in order to study how this in-context learning (ICL) capability emerges. In our setting, each example is sampled from a Markov chain drawn from a prior distribution over Markov chains. Transformers trained on this task form \emph{statistical induction heads} which compute accurate next-token probabilities given the bigram statistics of the context. During the course of training, models pass through multiple phases: after an initial stage in which predictions are uniform, they learn to sub-optimally predict using in-context single-token statistics (unigrams); then, there is a rapid phase transition to the correct in-context bigram solution. We conduct an empirical and theoretical investigation of this multi-phase process, showing how successful learning results from the interaction between the transformer's layers, and uncovering evidence that the presence of the simpler unigram solution may delay formation of the final bigram solution. We examine how learning is affected by varying the prior distribution over Markov chains, and consider the generalization of our in-context learning of Markov chains (ICL-MC) task to $n$-grams for $n > 2$.
title The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
topic Machine Learning
url https://arxiv.org/abs/2402.11004