The Quantization Model of Neural Scaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Michaud, Eric J., Liu, Ziming, Girit, Uzay, Tegmark, Max
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909072502030336
author Michaud, Eric J.
Liu, Ziming
Girit, Uzay
Tegmark, Max
author_facet Michaud, Eric J.
Liu, Ziming
Girit, Uzay
Tegmark, Max
contents We propose the Quantization Model of neural scaling laws, explaining both the observed power law dropoff of loss with model and data size, and also the sudden emergence of new capabilities with scale. We derive this model from what we call the Quantization Hypothesis, where network knowledge and skills are "quantized" into discrete chunks ($\textbf{quanta}$). We show that when quanta are learned in order of decreasing use frequency, then a power law in use frequencies explains observed power law scaling of loss. We validate this prediction on toy datasets, then study how scaling curves decompose for large language models. Using language model gradients, we automatically decompose model behavior into a diverse set of skills (quanta). We tentatively find that the frequency at which these quanta are used in the training distribution roughly follows a power law corresponding with the empirical scaling exponent for language models, a prediction of our theory.
format Preprint
id arxiv_https___arxiv_org_abs_2303_13506
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The Quantization Model of Neural Scaling
Michaud, Eric J.
Liu, Ziming
Girit, Uzay
Tegmark, Max
Machine Learning
Disordered Systems and Neural Networks
We propose the Quantization Model of neural scaling laws, explaining both the observed power law dropoff of loss with model and data size, and also the sudden emergence of new capabilities with scale. We derive this model from what we call the Quantization Hypothesis, where network knowledge and skills are "quantized" into discrete chunks ($\textbf{quanta}$). We show that when quanta are learned in order of decreasing use frequency, then a power law in use frequencies explains observed power law scaling of loss. We validate this prediction on toy datasets, then study how scaling curves decompose for large language models. Using language model gradients, we automatically decompose model behavior into a diverse set of skills (quanta). We tentatively find that the frequency at which these quanta are used in the training distribution roughly follows a power law corresponding with the empirical scaling exponent for language models, a prediction of our theory.
title The Quantization Model of Neural Scaling
topic Machine Learning
Disordered Systems and Neural Networks
url https://arxiv.org/abs/2303.13506