Efficient Parallelization Layouts for Large-Scale Distributed Model Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hagemann, Johannes, Weinbach, Samuel, Dobler, Konstantin, Schall, Maximilian, de Melo, Gerard
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913515144478720
author Hagemann, Johannes
Weinbach, Samuel
Dobler, Konstantin
Schall, Maximilian
de Melo, Gerard
author_facet Hagemann, Johannes
Weinbach, Samuel
Dobler, Konstantin
Schall, Maximilian
de Melo, Gerard
contents Efficiently training large language models requires parallelizing across hundreds of hardware accelerators and invoking various compute and memory optimizations. When combined, many of these strategies have complex interactions regarding the final training efficiency. Prior work tackling this problem did not have access to the latest set of optimizations, such as FlashAttention or sequence parallelism. In this work, we conduct a comprehensive ablation study of possible training configurations for large language models. We distill this large study into several key recommendations for the most efficient training. For instance, we find that using a micro-batch size of 1 usually enables the most efficient training layouts. Larger micro-batch sizes necessitate activation checkpointing or higher degrees of model parallelism and also lead to larger pipeline bubbles. Our most efficient configurations enable us to achieve state-of-the-art training efficiency results over a range of model sizes, most notably a Model FLOPs utilization of 70.5% when training a Llama 13B model.
format Preprint
id arxiv_https___arxiv_org_abs_2311_05610
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Efficient Parallelization Layouts for Large-Scale Distributed Model Training
Hagemann, Johannes
Weinbach, Samuel
Dobler, Konstantin
Schall, Maximilian
de Melo, Gerard
Machine Learning
Distributed, Parallel, and Cluster Computing
Efficiently training large language models requires parallelizing across hundreds of hardware accelerators and invoking various compute and memory optimizations. When combined, many of these strategies have complex interactions regarding the final training efficiency. Prior work tackling this problem did not have access to the latest set of optimizations, such as FlashAttention or sequence parallelism. In this work, we conduct a comprehensive ablation study of possible training configurations for large language models. We distill this large study into several key recommendations for the most efficient training. For instance, we find that using a micro-batch size of 1 usually enables the most efficient training layouts. Larger micro-batch sizes necessitate activation checkpointing or higher degrees of model parallelism and also lead to larger pipeline bubbles. Our most efficient configurations enable us to achieve state-of-the-art training efficiency results over a range of model sizes, most notably a Model FLOPs utilization of 70.5% when training a Llama 13B model.
title Efficient Parallelization Layouts for Large-Scale Distributed Model Training
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2311.05610