Improving Automatic Parallel Training via Balanced Memory Workload Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yujie, Jiang, Youhe, Miao, Xupeng, Fu, Fangcheng, Zhu, Shenhan, Nie, Xiaonan, Tu, Yaofeng, Cui, Bin
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912014829355008
author Wang, Yujie
Jiang, Youhe
Miao, Xupeng
Fu, Fangcheng
Zhu, Shenhan
Nie, Xiaonan
Tu, Yaofeng
Cui, Bin
author_facet Wang, Yujie
Jiang, Youhe
Miao, Xupeng
Fu, Fangcheng
Zhu, Shenhan
Nie, Xiaonan
Tu, Yaofeng
Cui, Bin
contents Transformer models have emerged as the leading approach for achieving state-of-the-art performance across various application domains, serving as the foundation for advanced large-scale deep learning (DL) models. However, efficiently training these models across multiple GPUs remains a complex challenge due to the abundance of parallelism options. Existing DL systems either require manual efforts to design distributed training plans or limit parallelism combinations to a constrained search space. In this paper, we present Galvatron-BMW, a novel system framework that integrates multiple prevalent parallelism dimensions and automatically identifies the most efficient hybrid parallelism strategy. To effectively navigate this vast search space, we employ a decision tree approach for decomposition and pruning based on intuitive insights. We further utilize a dynamic programming search algorithm to derive the optimal plan. Moreover, to improve resource utilization and enhance system efficiency, we propose a bi-objective optimization workflow that focuses on workload balance. Our evaluations on different Transformer models demonstrate the capabilities of Galvatron-BMW in automating distributed training under varying GPU memory constraints. Across all tested scenarios, Galvatron-BMW consistently achieves superior system throughput, surpassing previous approaches that rely on limited parallelism strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2307_02031
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Improving Automatic Parallel Training via Balanced Memory Workload Optimization
Wang, Yujie
Jiang, Youhe
Miao, Xupeng
Fu, Fangcheng
Zhu, Shenhan
Nie, Xiaonan
Tu, Yaofeng
Cui, Bin
Machine Learning
Databases
Distributed, Parallel, and Cluster Computing
Transformer models have emerged as the leading approach for achieving state-of-the-art performance across various application domains, serving as the foundation for advanced large-scale deep learning (DL) models. However, efficiently training these models across multiple GPUs remains a complex challenge due to the abundance of parallelism options. Existing DL systems either require manual efforts to design distributed training plans or limit parallelism combinations to a constrained search space. In this paper, we present Galvatron-BMW, a novel system framework that integrates multiple prevalent parallelism dimensions and automatically identifies the most efficient hybrid parallelism strategy. To effectively navigate this vast search space, we employ a decision tree approach for decomposition and pruning based on intuitive insights. We further utilize a dynamic programming search algorithm to derive the optimal plan. Moreover, to improve resource utilization and enhance system efficiency, we propose a bi-objective optimization workflow that focuses on workload balance. Our evaluations on different Transformer models demonstrate the capabilities of Galvatron-BMW in automating distributed training under varying GPU memory constraints. Across all tested scenarios, Galvatron-BMW consistently achieves superior system throughput, surpassing previous approaches that rely on limited parallelism strategies.
title Improving Automatic Parallel Training via Balanced Memory Workload Optimization
topic Machine Learning
Databases
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2307.02031