A Dynamical Model of Neural Scaling Laws

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bordelon, Blake, Atanasov, Alexander, Pehlevan, Cengiz
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929396375355392
author Bordelon, Blake
Atanasov, Alexander
Pehlevan, Cengiz
author_facet Bordelon, Blake
Atanasov, Alexander
Pehlevan, Cengiz
contents On a variety of tasks, the performance of neural networks predictably improves with training time, dataset size and model size across many orders of magnitude. This phenomenon is known as a neural scaling law. Of fundamental importance is the compute-optimal scaling law, which reports the performance as a function of units of compute when choosing model sizes optimally. We analyze a random feature model trained with gradient descent as a solvable model of network training and generalization. This reproduces many observations about neural scaling laws. First, our model makes a prediction about why the scaling of performance with training time and with model size have different power law exponents. Consequently, the theory predicts an asymmetric compute-optimal scaling rule where the number of training steps are increased faster than model parameters, consistent with recent empirical observations. Second, it has been observed that early in training, networks converge to their infinite-width dynamics at a rate $1/\textit{width}$ but at late time exhibit a rate $\textit{width}^{-c}$, where $c$ depends on the structure of the architecture and task. We show that our model exhibits this behavior. Lastly, our theory shows how the gap between training and test loss can gradually build up over time due to repeated reuse of data.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01092
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Dynamical Model of Neural Scaling Laws
Bordelon, Blake
Atanasov, Alexander
Pehlevan, Cengiz
Machine Learning
Disordered Systems and Neural Networks
On a variety of tasks, the performance of neural networks predictably improves with training time, dataset size and model size across many orders of magnitude. This phenomenon is known as a neural scaling law. Of fundamental importance is the compute-optimal scaling law, which reports the performance as a function of units of compute when choosing model sizes optimally. We analyze a random feature model trained with gradient descent as a solvable model of network training and generalization. This reproduces many observations about neural scaling laws. First, our model makes a prediction about why the scaling of performance with training time and with model size have different power law exponents. Consequently, the theory predicts an asymmetric compute-optimal scaling rule where the number of training steps are increased faster than model parameters, consistent with recent empirical observations. Second, it has been observed that early in training, networks converge to their infinite-width dynamics at a rate $1/\textit{width}$ but at late time exhibit a rate $\textit{width}^{-c}$, where $c$ depends on the structure of the architecture and task. We show that our model exhibits this behavior. Lastly, our theory shows how the gap between training and test loss can gradually build up over time due to repeated reuse of data.
title A Dynamical Model of Neural Scaling Laws
topic Machine Learning
Disordered Systems and Neural Networks
url https://arxiv.org/abs/2402.01092