An Optimized Error-controlled MPI Collective Framework Integrated with Lossy Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jiajun, Di, Sheng, Yu, Xiaodong, Zhai, Yujia, Zhang, Zhaorui, Liu, Jinyang, Lu, Xiaoyi, Raffenetti, Ken, Zhou, Hui, Zhao, Kai, Chen, Zizhong, Cappello, Franck, Guo, Yanfei, Thakur, Rajeev
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929213493215232
author Huang, Jiajun
Di, Sheng
Yu, Xiaodong
Zhai, Yujia
Zhang, Zhaorui
Liu, Jinyang
Lu, Xiaoyi
Raffenetti, Ken
Zhou, Hui
Zhao, Kai
Chen, Zizhong
Cappello, Franck
Guo, Yanfei
Thakur, Rajeev
author_facet Huang, Jiajun
Di, Sheng
Yu, Xiaodong
Zhai, Yujia
Zhang, Zhaorui
Liu, Jinyang
Lu, Xiaoyi
Raffenetti, Ken
Zhou, Hui
Zhao, Kai
Chen, Zizhong
Cappello, Franck
Guo, Yanfei
Thakur, Rajeev
contents With the ever-increasing computing power of supercomputers and the growing scale of scientific applications, the efficiency of MPI collective communications turns out to be a critical bottleneck in large-scale distributed and parallel processing. The large message size in MPI collectives is particularly concerning because it can significantly degrade the overall parallel performance. To address this issue, prior research simply applies the off-the-shelf fix-rate lossy compressors in the MPI collectives, leading to suboptimal performance, limited generalizability, and unbounded errors. In this paper, we propose a novel solution, called C-Coll, which leverages error-bounded lossy compression to significantly reduce the message size, resulting in a substantial reduction in communication cost. The key contributions are three-fold. (1) We develop two general, optimized lossy-compression-based frameworks for both types of MPI collectives (collective data movement as well as collective computation), based on their particular characteristics. Our framework not only reduces communication cost but also preserves data accuracy. (2) We customize SZx, an ultra-fast error-bounded lossy compressor, to meet the specific needs of collective communication. (3) We integrate C-Coll into multiple collectives, such as MPI_Allreduce, MPI_Scatter, and MPI_Bcast, and perform a comprehensive evaluation based on real-world scientific datasets. Experiments show that our solution outperforms the original MPI collectives as well as multiple baselines and related efforts by 1.8-2.7X.
format Preprint
id arxiv_https___arxiv_org_abs_2304_03890
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle An Optimized Error-controlled MPI Collective Framework Integrated with Lossy Compression
Huang, Jiajun
Di, Sheng
Yu, Xiaodong
Zhai, Yujia
Zhang, Zhaorui
Liu, Jinyang
Lu, Xiaoyi
Raffenetti, Ken
Zhou, Hui
Zhao, Kai
Chen, Zizhong
Cappello, Franck
Guo, Yanfei
Thakur, Rajeev
Distributed, Parallel, and Cluster Computing
With the ever-increasing computing power of supercomputers and the growing scale of scientific applications, the efficiency of MPI collective communications turns out to be a critical bottleneck in large-scale distributed and parallel processing. The large message size in MPI collectives is particularly concerning because it can significantly degrade the overall parallel performance. To address this issue, prior research simply applies the off-the-shelf fix-rate lossy compressors in the MPI collectives, leading to suboptimal performance, limited generalizability, and unbounded errors. In this paper, we propose a novel solution, called C-Coll, which leverages error-bounded lossy compression to significantly reduce the message size, resulting in a substantial reduction in communication cost. The key contributions are three-fold. (1) We develop two general, optimized lossy-compression-based frameworks for both types of MPI collectives (collective data movement as well as collective computation), based on their particular characteristics. Our framework not only reduces communication cost but also preserves data accuracy. (2) We customize SZx, an ultra-fast error-bounded lossy compressor, to meet the specific needs of collective communication. (3) We integrate C-Coll into multiple collectives, such as MPI_Allreduce, MPI_Scatter, and MPI_Bcast, and perform a comprehensive evaluation based on real-world scientific datasets. Experiments show that our solution outperforms the original MPI collectives as well as multiple baselines and related efforts by 1.8-2.7X.
title An Optimized Error-controlled MPI Collective Framework Integrated with Lossy Compression
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2304.03890