Deriving Neural Scaling Laws from the statistics of natural language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cagnetta, Francesco, Raventós, Allan, Ganguli, Surya, Wyart, Matthieu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918333301915648
author Cagnetta, Francesco
Raventós, Allan
Ganguli, Surya
Wyart, Matthieu
author_facet Cagnetta, Francesco
Raventós, Allan
Ganguli, Surya
Wyart, Matthieu
contents Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07488
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Deriving Neural Scaling Laws from the statistics of natural language
Cagnetta, Francesco
Raventós, Allan
Ganguli, Surya
Wyart, Matthieu
Machine Learning
Artificial Intelligence
Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.
title Deriving Neural Scaling Laws from the statistics of natural language
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.07488