Saved in:
Bibliographic Details
Main Author: Namashivayam, Naveen
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.24230
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912301849772032
author Namashivayam, Naveen
author_facet Namashivayam, Naveen
contents Compute nodes on modern heterogeneous supercomputing systems comprise CPUs, GPUs, and high-speed network interconnects (NICs). Parallelization is identified as a technique for effectively utilizing these systems to execute scalable simulation and deep learning workloads. The resulting inter-process communication from the distributed execution of these parallel workloads is one of the key factors contributing to its performance bottleneck. Most programming models and runtime systems enabling the communication requirements on these systems support GPU-aware communication schemes that move the GPU-attached communication buffers in the application directly from the GPU to the NIC without staging through the host memory. A CPU thread is required to orchestrate the communication operations even with support for such GPU-awareness. This survey discusses various available GPU-centric communication schemes that move the control path of the communication operations from the CPU to the GPU. This work presents the need for the new communication schemes, various GPU and NIC capabilities required to implement the schemes, and the potential use-cases addressed. Based on these discussions, challenges involved in supporting the exhibited GPU-centric communication schemes are discussed.
format Preprint
id arxiv_https___arxiv_org_abs_2503_24230
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GPU-centric Communication Schemes for HPC and ML Applications
Namashivayam, Naveen
Distributed, Parallel, and Cluster Computing
Machine Learning
C.2.0; C.4; C.5.1; D.1.3
Compute nodes on modern heterogeneous supercomputing systems comprise CPUs, GPUs, and high-speed network interconnects (NICs). Parallelization is identified as a technique for effectively utilizing these systems to execute scalable simulation and deep learning workloads. The resulting inter-process communication from the distributed execution of these parallel workloads is one of the key factors contributing to its performance bottleneck. Most programming models and runtime systems enabling the communication requirements on these systems support GPU-aware communication schemes that move the GPU-attached communication buffers in the application directly from the GPU to the NIC without staging through the host memory. A CPU thread is required to orchestrate the communication operations even with support for such GPU-awareness. This survey discusses various available GPU-centric communication schemes that move the control path of the communication operations from the CPU to the GPU. This work presents the need for the new communication schemes, various GPU and NIC capabilities required to implement the schemes, and the potential use-cases addressed. Based on these discussions, challenges involved in supporting the exhibited GPU-centric communication schemes are discussed.
title GPU-centric Communication Schemes for HPC and ML Applications
topic Distributed, Parallel, and Cluster Computing
Machine Learning
C.2.0; C.4; C.5.1; D.1.3
url https://arxiv.org/abs/2503.24230