Neural Scaling Laws Rooted in the Data Distribution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Brill, Ari
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913607463206912
author Brill, Ari
author_facet Brill, Ari
contents Deep neural networks exhibit empirical neural scaling laws, with error decreasing as a power law with increasing model or data size, across a wide variety of architectures, tasks, and datasets. This universality suggests that scaling laws may result from general properties of natural learning tasks. We develop a mathematical model intended to describe natural datasets using percolation theory. Two distinct criticality regimes emerge, each yielding optimal power-law neural scaling laws. These regimes, corresponding to power-law-distributed discrete subtasks and a dominant data manifold, can be associated with previously proposed theories of neural scaling, thereby grounding and unifying prior works. We test the theory by training regression models on toy datasets derived from percolation theory simulations. We suggest directions for quantitatively predicting language model scaling.
format Preprint
id arxiv_https___arxiv_org_abs_2412_07942
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Neural Scaling Laws Rooted in the Data Distribution
Brill, Ari
Machine Learning
Disordered Systems and Neural Networks
Deep neural networks exhibit empirical neural scaling laws, with error decreasing as a power law with increasing model or data size, across a wide variety of architectures, tasks, and datasets. This universality suggests that scaling laws may result from general properties of natural learning tasks. We develop a mathematical model intended to describe natural datasets using percolation theory. Two distinct criticality regimes emerge, each yielding optimal power-law neural scaling laws. These regimes, corresponding to power-law-distributed discrete subtasks and a dominant data manifold, can be associated with previously proposed theories of neural scaling, thereby grounding and unifying prior works. We test the theory by training regression models on toy datasets derived from percolation theory simulations. We suggest directions for quantitatively predicting language model scaling.
title Neural Scaling Laws Rooted in the Data Distribution
topic Machine Learning
Disordered Systems and Neural Networks
url https://arxiv.org/abs/2412.07942