Next-token pretraining implies in-context learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Riechers, Paul M., Bigelow, Henry R., Alt, Eric A., Shai, Adam
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911052408553472
author Riechers, Paul M.
Bigelow, Henry R.
Alt, Eric A.
Shai, Adam
author_facet Riechers, Paul M.
Bigelow, Henry R.
Alt, Eric A.
Shai, Adam
contents We argue that in-context learning (ICL) predictably arises from standard self-supervised next-token pretraining, rather than being an exotic emergent property. This work establishes the foundational principles of this emergence by focusing on in-distribution ICL, demonstrating how models necessarily adapt to context when trained on token sequences, especially from non-ergodic sources. Our information-theoretic framework precisely predicts these in-distribution ICL dynamics (i.e., context-dependent loss reduction). We verify this with experiments using synthetic datasets of differing types of correlational structure, reproducing characteristic phenomena like phase transitions in training loss for induction head formation and power-law scaling of in-context loss. We further show that a model's in-context performance on any task is mathematically coupled to the ensemble of tasks seen in pretraining, offering a fundamental explanation, grounded in architecture- and modality-independent principles, for such inference-time learning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Next-token pretraining implies in-context learning
Riechers, Paul M.
Bigelow, Henry R.
Alt, Eric A.
Shai, Adam
Machine Learning
Artificial Intelligence
We argue that in-context learning (ICL) predictably arises from standard self-supervised next-token pretraining, rather than being an exotic emergent property. This work establishes the foundational principles of this emergence by focusing on in-distribution ICL, demonstrating how models necessarily adapt to context when trained on token sequences, especially from non-ergodic sources. Our information-theoretic framework precisely predicts these in-distribution ICL dynamics (i.e., context-dependent loss reduction). We verify this with experiments using synthetic datasets of differing types of correlational structure, reproducing characteristic phenomena like phase transitions in training loss for induction head formation and power-law scaling of in-context loss. We further show that a model's in-context performance on any task is mathematically coupled to the ensemble of tasks seen in pretraining, offering a fundamental explanation, grounded in architecture- and modality-independent principles, for such inference-time learning.
title Next-token pretraining implies in-context learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.18373