On Optimizing the Communication of Model Parallelism

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Yonghao, Zhao, Hexu, Zheng, Lianmin, Li, Zhuohan, Xing, Eric P., Ho, Qirong, Gonzalez, Joseph E., Stoica, Ion, Zhang, Hao
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913470128062464
author Zhuang, Yonghao
Zhao, Hexu
Zheng, Lianmin
Li, Zhuohan
Xing, Eric P.
Ho, Qirong
Gonzalez, Joseph E.
Stoica, Ion
Zhang, Hao
author_facet Zhuang, Yonghao
Zhao, Hexu
Zheng, Lianmin
Li, Zhuohan
Xing, Eric P.
Ho, Qirong
Gonzalez, Joseph E.
Stoica, Ion
Zhang, Hao
contents We study a novel and important communication pattern in large-scale model-parallel deep learning (DL), which we call cross-mesh resharding. This pattern emerges when the two paradigms of model parallelism - intra-operator and inter-operator parallelism - are combined to support large models on large clusters. In cross-mesh resharding, a sharded tensor needs to be sent from a source device mesh to a destination device mesh, on which the tensor may be distributed with the same or different layouts. We formalize this as a many-to-many multicast communication problem, and show that existing approaches either are sub-optimal or do not generalize to different network topologies or tensor layouts, which result from different model architectures and parallelism strategies. We then propose two contributions to address cross-mesh resharding: an efficient broadcast-based communication system, and an "overlapping-friendly" pipeline schedule. On microbenchmarks, our overall system outperforms existing ones by up to 10x across various tensor and mesh layouts. On end-to-end training of two large models, GPT-3 and U-Transformer, we improve throughput by 10% and 50%, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2211_05322
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle On Optimizing the Communication of Model Parallelism
Zhuang, Yonghao
Zhao, Hexu
Zheng, Lianmin
Li, Zhuohan
Xing, Eric P.
Ho, Qirong
Gonzalez, Joseph E.
Stoica, Ion
Zhang, Hao
Machine Learning
Distributed, Parallel, and Cluster Computing
We study a novel and important communication pattern in large-scale model-parallel deep learning (DL), which we call cross-mesh resharding. This pattern emerges when the two paradigms of model parallelism - intra-operator and inter-operator parallelism - are combined to support large models on large clusters. In cross-mesh resharding, a sharded tensor needs to be sent from a source device mesh to a destination device mesh, on which the tensor may be distributed with the same or different layouts. We formalize this as a many-to-many multicast communication problem, and show that existing approaches either are sub-optimal or do not generalize to different network topologies or tensor layouts, which result from different model architectures and parallelism strategies. We then propose two contributions to address cross-mesh resharding: an efficient broadcast-based communication system, and an "overlapping-friendly" pipeline schedule. On microbenchmarks, our overall system outperforms existing ones by up to 10x across various tensor and mesh layouts. On end-to-end training of two large models, GPT-3 and U-Transformer, we improve throughput by 10% and 50%, respectively.
title On Optimizing the Communication of Model Parallelism
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2211.05322