Predicting the Emergence of Induction Heads in Language Model Pretraining

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aoyama, Tatsuya, Wilcox, Ethan Gotlieb, Schneider, Nathan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917256369274880
author Aoyama, Tatsuya
Wilcox, Ethan Gotlieb
Schneider, Nathan
author_facet Aoyama, Tatsuya
Wilcox, Ethan Gotlieb
Schneider, Nathan
contents Specialized attention heads dubbed induction heads (IHs) have been argued to underlie the remarkable in-context learning capabilities of modern language models; yet, a precise characterization of their emergence, especially in the context of language modeling, remains wanting. In this study, we investigate the relationship between statistical properties of the training data and IH formation in both natural and synthetic training data settings. We show that: (1) A simple equation combining batch size and context size predicts the point at which IHs form and that this emergence point is agnostic to model size; (2) Surface bigram repetition frequency and reliability strongly affect the formation of IHs, and we find an effective Pareto frontier in terms of these two values; (3) local dependency with high bigram repetition frequency and reliability is sufficient for IH formation, but when the frequency and reliability are low, categoriality and the shape of the marginal distribution matter.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16893
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Predicting the Emergence of Induction Heads in Language Model Pretraining
Aoyama, Tatsuya
Wilcox, Ethan Gotlieb
Schneider, Nathan
Computation and Language
Specialized attention heads dubbed induction heads (IHs) have been argued to underlie the remarkable in-context learning capabilities of modern language models; yet, a precise characterization of their emergence, especially in the context of language modeling, remains wanting. In this study, we investigate the relationship between statistical properties of the training data and IH formation in both natural and synthetic training data settings. We show that: (1) A simple equation combining batch size and context size predicts the point at which IHs form and that this emergence point is agnostic to model size; (2) Surface bigram repetition frequency and reliability strongly affect the formation of IHs, and we find an effective Pareto frontier in terms of these two values; (3) local dependency with high bigram repetition frequency and reliability is sufficient for IH formation, but when the frequency and reliability are low, categoriality and the shape of the marginal distribution matter.
title Predicting the Emergence of Induction Heads in Language Model Pretraining
topic Computation and Language
url https://arxiv.org/abs/2511.16893