Time Matters: Scaling Laws for Any Budget

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Inbar, Itay, Sernau, Luke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909361309220864
author Inbar, Itay
Sernau, Luke
author_facet Inbar, Itay
Sernau, Luke
contents A primary cost driver for training large models is wall-clock training time. We show that popular time estimates based on FLOPs are poor estimates, and construct a more accurate proxy based on memory copies. This allows us to accurately estimate the training speed of a transformer model from its hyperparameters. Combined with a scaling law curve like Chinchilla, this allows us to accurately predict the final loss of a model from a simple equation. We show that this expression is accurate across a wide range of model hyperparameter values, enabling us to analytically make architectural decisions and train models more efficiently. Crucially, this analysis predicts that in contrast to existing literature, models should be wider rather than deeper, as the benefits of speed outweigh the benefits of depth.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18922
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Time Matters: Scaling Laws for Any Budget
Inbar, Itay
Sernau, Luke
Machine Learning
Artificial Intelligence
A primary cost driver for training large models is wall-clock training time. We show that popular time estimates based on FLOPs are poor estimates, and construct a more accurate proxy based on memory copies. This allows us to accurately estimate the training speed of a transformer model from its hyperparameters. Combined with a scaling law curve like Chinchilla, this allows us to accurately predict the final loss of a model from a simple equation. We show that this expression is accurate across a wide range of model hyperparameter values, enabling us to analytically make architectural decisions and train models more efficiently. Crucially, this analysis predicts that in contrast to existing literature, models should be wider rather than deeper, as the benefits of speed outweigh the benefits of depth.
title Time Matters: Scaling Laws for Any Budget
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.18922