On Optimizing the Communication of Model Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913470128062464 |
|---|---|
| author | Zhuang, Yonghao Zhao, Hexu Zheng, Lianmin Li, Zhuohan Xing, Eric P. Ho, Qirong Gonzalez, Joseph E. Stoica, Ion Zhang, Hao |
| author_facet | Zhuang, Yonghao Zhao, Hexu Zheng, Lianmin Li, Zhuohan Xing, Eric P. Ho, Qirong Gonzalez, Joseph E. Stoica, Ion Zhang, Hao |
| contents | We study a novel and important communication pattern in large-scale model-parallel deep learning (DL), which we call cross-mesh resharding. This pattern emerges when the two paradigms of model parallelism - intra-operator and inter-operator parallelism - are combined to support large models on large clusters. In cross-mesh resharding, a sharded tensor needs to be sent from a source device mesh to a destination device mesh, on which the tensor may be distributed with the same or different layouts. We formalize this as a many-to-many multicast communication problem, and show that existing approaches either are sub-optimal or do not generalize to different network topologies or tensor layouts, which result from different model architectures and parallelism strategies. We then propose two contributions to address cross-mesh resharding: an efficient broadcast-based communication system, and an "overlapping-friendly" pipeline schedule. On microbenchmarks, our overall system outperforms existing ones by up to 10x across various tensor and mesh layouts. On end-to-end training of two large models, GPT-3 and U-Transformer, we improve throughput by 10% and 50%, respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2211_05322 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | On Optimizing the Communication of Model Parallelism Zhuang, Yonghao Zhao, Hexu Zheng, Lianmin Li, Zhuohan Xing, Eric P. Ho, Qirong Gonzalez, Joseph E. Stoica, Ion Zhang, Hao Machine Learning Distributed, Parallel, and Cluster Computing We study a novel and important communication pattern in large-scale model-parallel deep learning (DL), which we call cross-mesh resharding. This pattern emerges when the two paradigms of model parallelism - intra-operator and inter-operator parallelism - are combined to support large models on large clusters. In cross-mesh resharding, a sharded tensor needs to be sent from a source device mesh to a destination device mesh, on which the tensor may be distributed with the same or different layouts. We formalize this as a many-to-many multicast communication problem, and show that existing approaches either are sub-optimal or do not generalize to different network topologies or tensor layouts, which result from different model architectures and parallelism strategies. We then propose two contributions to address cross-mesh resharding: an efficient broadcast-based communication system, and an "overlapping-friendly" pipeline schedule. On microbenchmarks, our overall system outperforms existing ones by up to 10x across various tensor and mesh layouts. On end-to-end training of two large models, GPT-3 and U-Transformer, we improve throughput by 10% and 50%, respectively. |
| title | On Optimizing the Communication of Model Parallelism |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2211.05322 |