Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, WenZheng, Hu, Yang, Shi, Jing, Bai, Xiaoying
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913476882989056
author Zhang, WenZheng
Hu, Yang
Shi, Jing
Bai, Xiaoying
author_facet Zhang, WenZheng
Hu, Yang
Shi, Jing
Bai, Xiaoying
contents Scaling Deep Neural Networks (DNNs) requires significant computational resources in terms of GPU quantity and compute capacity. In practice, there usually exists a large number of heterogeneous GPU devices due to the rapid release cycle of GPU products. It is highly needed to efficiently and economically harness the power of heterogeneous GPUs, so that it can meet the requirements of DNN research and development. The paper introduces Poplar, a distributed training system that extends Zero Redundancy Optimizer (ZeRO) with heterogeneous-aware capabilities. We explore a broader spectrum of GPU heterogeneity, including compute capability, memory capacity, quantity and a combination of them. In order to achieve high computational efficiency across all heterogeneous conditions, Poplar conducts fine-grained measurements of GPUs in each ZeRO stage. We propose a novel batch allocation method and a search algorithm to optimize the utilization of heterogeneous GPUs clusters. Furthermore, Poplar implements fully automated parallelism, eliminating the need for deploying heterogeneous hardware and finding suitable batch size. Extensive experiments on three heterogeneous clusters, comprising six different types of GPUs, demonstrate that Poplar achieves a training throughput improvement of 1.02-3.92x over current state-of-the-art heterogeneous training systems.
format Preprint
id arxiv_https___arxiv_org_abs_2408_12596
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
Zhang, WenZheng
Hu, Yang
Shi, Jing
Bai, Xiaoying
Distributed, Parallel, and Cluster Computing
Scaling Deep Neural Networks (DNNs) requires significant computational resources in terms of GPU quantity and compute capacity. In practice, there usually exists a large number of heterogeneous GPU devices due to the rapid release cycle of GPU products. It is highly needed to efficiently and economically harness the power of heterogeneous GPUs, so that it can meet the requirements of DNN research and development. The paper introduces Poplar, a distributed training system that extends Zero Redundancy Optimizer (ZeRO) with heterogeneous-aware capabilities. We explore a broader spectrum of GPU heterogeneity, including compute capability, memory capacity, quantity and a combination of them. In order to achieve high computational efficiency across all heterogeneous conditions, Poplar conducts fine-grained measurements of GPUs in each ZeRO stage. We propose a novel batch allocation method and a search algorithm to optimize the utilization of heterogeneous GPUs clusters. Furthermore, Poplar implements fully automated parallelism, eliminating the need for deploying heterogeneous hardware and finding suitable batch size. Extensive experiments on three heterogeneous clusters, comprising six different types of GPUs, demonstrate that Poplar achieves a training throughput improvement of 1.02-3.92x over current state-of-the-art heterogeneous training systems.
title Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2408.12596