Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pan, Xinglin, Lin, Wenxiang, Shi, Shaohuai, Chu, Xiaowen, Sun, Weinong, Li, Bo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916310240198656
author Pan, Xinglin
Lin, Wenxiang
Shi, Shaohuai
Chu, Xiaowen
Sun, Weinong
Li, Bo
author_facet Pan, Xinglin
Lin, Wenxiang
Shi, Shaohuai
Chu, Xiaowen
Sun, Weinong
Li, Bo
contents Sparsely-activated Mixture-of-Expert (MoE) layers have found practical applications in enlarging the model size of large-scale foundation models, with only a sub-linear increase in computation demands. Despite the wide adoption of hybrid parallel paradigms like model parallelism, expert parallelism, and expert-sharding parallelism (i.e., MP+EP+ESP) to support MoE model training on GPU clusters, the training efficiency is hindered by communication costs introduced by these parallel paradigms. To address this limitation, we propose Parm, a system that accelerates MP+EP+ESP training by designing two dedicated schedules for placing communication tasks. The proposed schedules eliminate redundant computations and communications and enable overlaps between intra-node and inter-node communications, ultimately reducing the overall training time. As the two schedules are not mutually exclusive, we provide comprehensive theoretical analyses and derive an automatic and accurate solution to determine which schedule should be applied in different scenarios. Experimental results on an 8-GPU server and a 32-GPU cluster demonstrate that Parm outperforms the state-of-the-art MoE training system, DeepSpeed-MoE, achieving 1.13$\times$ to 5.77$\times$ speedup on 1296 manually configured MoE layers and approximately 3$\times$ improvement on two real-world MoE models based on BERT and GPT-2.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00599
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
Pan, Xinglin
Lin, Wenxiang
Shi, Shaohuai
Chu, Xiaowen
Sun, Weinong
Li, Bo
Distributed, Parallel, and Cluster Computing
Machine Learning
Sparsely-activated Mixture-of-Expert (MoE) layers have found practical applications in enlarging the model size of large-scale foundation models, with only a sub-linear increase in computation demands. Despite the wide adoption of hybrid parallel paradigms like model parallelism, expert parallelism, and expert-sharding parallelism (i.e., MP+EP+ESP) to support MoE model training on GPU clusters, the training efficiency is hindered by communication costs introduced by these parallel paradigms. To address this limitation, we propose Parm, a system that accelerates MP+EP+ESP training by designing two dedicated schedules for placing communication tasks. The proposed schedules eliminate redundant computations and communications and enable overlaps between intra-node and inter-node communications, ultimately reducing the overall training time. As the two schedules are not mutually exclusive, we provide comprehensive theoretical analyses and derive an automatic and accurate solution to determine which schedule should be applied in different scenarios. Experimental results on an 8-GPU server and a 32-GPU cluster demonstrate that Parm outperforms the state-of-the-art MoE training system, DeepSpeed-MoE, achieving 1.13$\times$ to 5.77$\times$ speedup on 1296 manually configured MoE layers and approximately 3$\times$ improvement on two real-world MoE models based on BERT and GPT-2.
title Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2407.00599