Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Shengwei, Lai, Zhiquan, Li, Dongsheng, Hao, Yanqi, Liu, Weijie, Ge, Keshi, Deng, Xiaoge, Lu, Kai
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916815878225920
author Li, Shengwei
Lai, Zhiquan
Li, Dongsheng
Hao, Yanqi
Liu, Weijie
Ge, Keshi
Deng, Xiaoge
Lu, Kai
author_facet Li, Shengwei
Lai, Zhiquan
Li, Dongsheng
Hao, Yanqi
Liu, Weijie
Ge, Keshi
Deng, Xiaoge
Lu, Kai
contents Deep learning is experiencing a rise in large-scale models. Training large-scale models is costly, prompting researchers to train large-scale models on commodity servers that more researchers can access. The massive number of parameters necessitates the use of model parallelism training methods. Existing studies focus on training with pipeline model parallelism. However, the tensor model parallelism (TMP) is inevitable when the model size keeps increasing, where frequent data-dependent communication and computation operations significantly reduce the training efficiency. In this paper, we present Oases, an automated TMP method with overlapped communication to accelerate large-scale model training on commodity servers. Oases proposes a fine-grained training operation schedule to maximize overlapping communication and computation that have data dependence. Additionally, we design the Oases planner that searches for the best model parameter partition strategy of TMP to achieve further accelerations. Unlike existing methods, Oases planner is tailored to model the cost of overlapped communication-computation operations. We evaluate Oases on various model settings and two commodity clusters, and compare Oases to four state-of-the-art implementations. Experimental results show that Oases achieves speedups of 1.01--1.48\(\times\) over the fastest baseline, and speedups of up to 1.95\(\times\) over Megatron.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16121
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
Li, Shengwei
Lai, Zhiquan
Li, Dongsheng
Hao, Yanqi
Liu, Weijie
Ge, Keshi
Deng, Xiaoge
Lu, Kai
Distributed, Parallel, and Cluster Computing
Deep learning is experiencing a rise in large-scale models. Training large-scale models is costly, prompting researchers to train large-scale models on commodity servers that more researchers can access. The massive number of parameters necessitates the use of model parallelism training methods. Existing studies focus on training with pipeline model parallelism. However, the tensor model parallelism (TMP) is inevitable when the model size keeps increasing, where frequent data-dependent communication and computation operations significantly reduce the training efficiency. In this paper, we present Oases, an automated TMP method with overlapped communication to accelerate large-scale model training on commodity servers. Oases proposes a fine-grained training operation schedule to maximize overlapping communication and computation that have data dependence. Additionally, we design the Oases planner that searches for the best model parameter partition strategy of TMP to achieve further accelerations. Unlike existing methods, Oases planner is tailored to model the cost of overlapped communication-computation operations. We evaluate Oases on various model settings and two commodity clusters, and compare Oases to four state-of-the-art implementations. Experimental results show that Oases achieves speedups of 1.01--1.48\(\times\) over the fastest baseline, and speedups of up to 1.95\(\times\) over Megatron.
title Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2305.16121