gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917658823229440 |
|---|---|
| author | Huang, Jiajun Di, Sheng Yu, Xiaodong Zhai, Yujia Liu, Jinyang Huang, Yafan Raffenetti, Ken Zhou, Hui Zhao, Kai Lu, Xiaoyi Chen, Zizhong Cappello, Franck Guo, Yanfei Thakur, Rajeev |
| author_facet | Huang, Jiajun Di, Sheng Yu, Xiaodong Zhai, Yujia Liu, Jinyang Huang, Yafan Raffenetti, Ken Zhou, Hui Zhao, Kai Lu, Xiaoyi Chen, Zizhong Cappello, Franck Guo, Yanfei Thakur, Rajeev |
| contents | GPU-aware collective communication has become a major bottleneck for modern computing platforms as GPU computing power rapidly rises. A traditional approach is to directly integrate lossy compression into GPU-aware collectives, which can lead to serious performance issues such as underutilized GPU devices and uncontrolled data distortion. In order to address these issues, in this paper, we propose gZCCL, a first-ever general framework that designs and optimizes GPU-aware, compression-enabled collectives with an accuracy-aware design to control error propagation. To validate our framework, we evaluate the performance on up to 512 NVIDIA A100 GPUs with real-world applications and datasets. Experimental results demonstrate that our gZCCL-accelerated collectives, including both collective computation (Allreduce) and collective data movement (Scatter), can outperform NCCL as well as Cray MPI by up to 4.5X and 28.7X, respectively. Furthermore, our accuracy evaluation with an image-stacking application confirms the high reconstructed data quality of our accuracy-aware framework. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2308_05199 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters Huang, Jiajun Di, Sheng Yu, Xiaodong Zhai, Yujia Liu, Jinyang Huang, Yafan Raffenetti, Ken Zhou, Hui Zhao, Kai Lu, Xiaoyi Chen, Zizhong Cappello, Franck Guo, Yanfei Thakur, Rajeev Distributed, Parallel, and Cluster Computing GPU-aware collective communication has become a major bottleneck for modern computing platforms as GPU computing power rapidly rises. A traditional approach is to directly integrate lossy compression into GPU-aware collectives, which can lead to serious performance issues such as underutilized GPU devices and uncontrolled data distortion. In order to address these issues, in this paper, we propose gZCCL, a first-ever general framework that designs and optimizes GPU-aware, compression-enabled collectives with an accuracy-aware design to control error propagation. To validate our framework, we evaluate the performance on up to 512 NVIDIA A100 GPUs with real-world applications and datasets. Experimental results demonstrate that our gZCCL-accelerated collectives, including both collective computation (Allreduce) and collective data movement (Scatter), can outperform NCCL as well as Cray MPI by up to 4.5X and 28.7X, respectively. Furthermore, our accuracy evaluation with an image-stacking application confirms the high reconstructed data quality of our accuracy-aware framework. |
| title | gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2308.05199 |