Data movement limits to frontier model training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Erdil, Ege, Schneider-Joseph, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913577110077440
author Erdil, Ege
Schneider-Joseph, David
author_facet Erdil, Ege
Schneider-Joseph, David
contents We present a theoretical model of distributed training, and use it to analyze how far dense and sparse training runs can be scaled. Under our baseline assumptions, given a three month training duration, data movement bottlenecks begin to significantly lower hardware utilization for training runs exceeding about $10^{28}$ FLOP, two orders of magnitude above the largest training run to date, suggesting the arrival of fundamental barriers to scaling in three years given recent rates of growth. A training run exceeding about $10^{31}$ FLOP is infeasible even at low utilization. However, more aggressive batch size scaling and/or shorter and fatter model shapes, if achievable, have the potential to permit much larger training runs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01137
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data movement limits to frontier model training
Erdil, Ege
Schneider-Joseph, David
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
We present a theoretical model of distributed training, and use it to analyze how far dense and sparse training runs can be scaled. Under our baseline assumptions, given a three month training duration, data movement bottlenecks begin to significantly lower hardware utilization for training runs exceeding about $10^{28}$ FLOP, two orders of magnitude above the largest training run to date, suggesting the arrival of fundamental barriers to scaling in three years given recent rates of growth. A training run exceeding about $10^{31}$ FLOP is infeasible even at low utilization. However, more aggressive batch size scaling and/or shorter and fatter model shapes, if achievable, have the potential to permit much larger training runs.
title Data movement limits to frontier model training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.01137