Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fernandez, Jared, Wehrstedt, Luca, Shamis, Leonid, Elhoushi, Mostafa, Saladi, Kalyan, Bisk, Yonatan, Strubell, Emma, Kahn, Jacob
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916687079538688
author Fernandez, Jared
Wehrstedt, Luca
Shamis, Leonid
Elhoushi, Mostafa
Saladi, Kalyan
Bisk, Yonatan
Strubell, Emma
Kahn, Jacob
author_facet Fernandez, Jared
Wehrstedt, Luca
Shamis, Leonid
Elhoushi, Mostafa
Saladi, Kalyan
Bisk, Yonatan
Strubell, Emma
Kahn, Jacob
contents Dramatic increases in the capabilities of neural network models in recent years are driven by scaling model size, training data, and corresponding computational resources. To develop the exceedingly large networks required in modern applications, such as large language models (LLMs), model training is distributed across tens of thousands of hardware accelerators (e.g. GPUs), requiring orchestration of computation and communication across large computing clusters. In this work, we demonstrate that careful consideration of hardware configuration and parallelization strategy is critical for effective (i.e. compute- and cost-efficient) scaling of model size, training data, and total computation. We conduct an extensive empirical study of the performance of large-scale LLM training workloads across model size, hardware configurations, and distributed parallelization strategies. We demonstrate that: (1) beyond certain scales, overhead incurred from certain distributed communication strategies leads parallelization strategies previously thought to be sub-optimal in fact become preferable; and (2) scaling the total number of accelerators for large model training quickly yields diminishing returns even when hardware and parallelization strategies are properly optimized, implying poor marginal performance per additional unit of power or GPU-hour.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13055
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
Fernandez, Jared
Wehrstedt, Luca
Shamis, Leonid
Elhoushi, Mostafa
Saladi, Kalyan
Bisk, Yonatan
Strubell, Emma
Kahn, Jacob
Machine Learning
Distributed, Parallel, and Cluster Computing
Dramatic increases in the capabilities of neural network models in recent years are driven by scaling model size, training data, and corresponding computational resources. To develop the exceedingly large networks required in modern applications, such as large language models (LLMs), model training is distributed across tens of thousands of hardware accelerators (e.g. GPUs), requiring orchestration of computation and communication across large computing clusters. In this work, we demonstrate that careful consideration of hardware configuration and parallelization strategy is critical for effective (i.e. compute- and cost-efficient) scaling of model size, training data, and total computation. We conduct an extensive empirical study of the performance of large-scale LLM training workloads across model size, hardware configurations, and distributed parallelization strategies. We demonstrate that: (1) beyond certain scales, overhead incurred from certain distributed communication strategies leads parallelization strategies previously thought to be sub-optimal in fact become preferable; and (2) scaling the total number of accelerators for large model training quickly yields diminishing returns even when hardware and parallelization strategies are properly optimized, implying poor marginal performance per additional unit of power or GPU-hour.
title Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2411.13055