Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Haoyang, Fu, Fangcheng, Ge, Hao, Lin, Sheng, Wang, Xuanyu, Niu, Jiawen, Wang, Yujie, Zhang, Hailin, Nie, Xiaonan, Cui, Bin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915274610966528
author Li, Haoyang
Fu, Fangcheng
Ge, Hao
Lin, Sheng
Wang, Xuanyu
Niu, Jiawen
Wang, Yujie
Zhang, Hailin
Nie, Xiaonan
Cui, Bin
author_facet Li, Haoyang
Fu, Fangcheng
Ge, Hao
Lin, Sheng
Wang, Xuanyu
Niu, Jiawen
Wang, Yujie
Zhang, Hailin
Nie, Xiaonan
Cui, Bin
contents As the scale of models and training data continues to grow, there is an expanding reliance on more GPUs to train large-scale models, which inevitably increases the likelihood of encountering dynamic stragglers that some devices lag behind in performance occasionally. However, hybrid parallel training, one of the de facto paradigms to train large models, is typically sensitive to the stragglers. This paper presents Malleus, a straggler-resilient hybrid parallel training framework for large-scale models. Malleus quantifies the stragglers at the nuanced, per-GPU granularity during training, and develops a novel planning algorithm to deduce the optimal parallelization of GPU devices, pipeline stages, model layers, and training data, maximizing training efficiency when stragglers exist. In addition, once a shift in the straggler situation is detected, Malleus adaptively adjusts the parallelization via a re-planning process, and seamlessly and efficiently migrates the model states on the fly, without sacrificing the stability of the training tasks. Empirical results on large language models with up to 110B parameters show that Malleus consistently outperforms existing parallel training frameworks under various straggler situations, delivering on average 2.63-5.28 times of efficiency improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13333
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
Li, Haoyang
Fu, Fangcheng
Ge, Hao
Lin, Sheng
Wang, Xuanyu
Niu, Jiawen
Wang, Yujie
Zhang, Hailin
Nie, Xiaonan
Cui, Bin
Distributed, Parallel, and Cluster Computing
As the scale of models and training data continues to grow, there is an expanding reliance on more GPUs to train large-scale models, which inevitably increases the likelihood of encountering dynamic stragglers that some devices lag behind in performance occasionally. However, hybrid parallel training, one of the de facto paradigms to train large models, is typically sensitive to the stragglers. This paper presents Malleus, a straggler-resilient hybrid parallel training framework for large-scale models. Malleus quantifies the stragglers at the nuanced, per-GPU granularity during training, and develops a novel planning algorithm to deduce the optimal parallelization of GPU devices, pipeline stages, model layers, and training data, maximizing training efficiency when stragglers exist. In addition, once a shift in the straggler situation is detected, Malleus adaptively adjusts the parallelization via a re-planning process, and seamlessly and efficiently migrates the model states on the fly, without sacrificing the stability of the training tasks. Empirical results on large language models with up to 110B parameters show that Malleus consistently outperforms existing parallel training frameworks under various straggler situations, delivering on average 2.63-5.28 times of efficiency improvement.
title Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2410.13333