Scaling Laws for Precision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Tanishq, Ankner, Zachary, Spector, Benjamin F., Bordelon, Blake, Muennighoff, Niklas, Paul, Mansheej, Pehlevan, Cengiz, Ré, Christopher, Raghunathan, Aditi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916500401553408
author Kumar, Tanishq
Ankner, Zachary
Spector, Benjamin F.
Bordelon, Blake
Muennighoff, Niklas
Paul, Mansheej
Pehlevan, Cengiz
Ré, Christopher
Raghunathan, Aditi
author_facet Kumar, Tanishq
Ankner, Zachary
Spector, Benjamin F.
Bordelon, Blake
Muennighoff, Niklas
Paul, Mansheej
Pehlevan, Cengiz
Ré, Christopher
Raghunathan, Aditi
contents Low precision training and inference affect both the quality and cost of language models, but current scaling laws do not account for this. In this work, we devise "precision-aware" scaling laws for both training and inference. We propose that training in lower precision reduces the model's "effective parameter count," allowing us to predict the additional loss incurred from training in low precision and post-train quantization. For inference, we find that the degradation introduced by post-training quantization increases as models are trained on more data, eventually making additional pretraining data actively harmful. For training, our scaling laws allow us to predict the loss of a model with different parts in different precisions, and suggest that training larger models in lower precision may be compute optimal. We unify the scaling laws for post and pretraining quantization to arrive at a single functional form that predicts degradation from training and inference in varied precisions. We fit on over 465 pretraining runs and validate our predictions on model sizes up to 1.7B parameters trained on up to 26B tokens.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04330
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scaling Laws for Precision
Kumar, Tanishq
Ankner, Zachary
Spector, Benjamin F.
Bordelon, Blake
Muennighoff, Niklas
Paul, Mansheej
Pehlevan, Cengiz
Ré, Christopher
Raghunathan, Aditi
Machine Learning
Computation and Language
Low precision training and inference affect both the quality and cost of language models, but current scaling laws do not account for this. In this work, we devise "precision-aware" scaling laws for both training and inference. We propose that training in lower precision reduces the model's "effective parameter count," allowing us to predict the additional loss incurred from training in low precision and post-train quantization. For inference, we find that the degradation introduced by post-training quantization increases as models are trained on more data, eventually making additional pretraining data actively harmful. For training, our scaling laws allow us to predict the loss of a model with different parts in different precisions, and suggest that training larger models in lower precision may be compute optimal. We unify the scaling laws for post and pretraining quantization to arrive at a single functional form that predicts degradation from training and inference in varied precisions. We fit on over 465 pretraining runs and validate our predictions on model sizes up to 1.7B parameters trained on up to 26B tokens.
title Scaling Laws for Precision
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2411.04330