RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Brock, Benjamin, Buluç, Aydın, Yelick, Katherine
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916274097881088
author Brock, Benjamin
Buluç, Aydın
Yelick, Katherine
author_facet Brock, Benjamin
Buluç, Aydın
Yelick, Katherine
contents Sparse matrix multiplication is an important kernel for large-scale graph processing and other data-intensive applications. In this paper, we implement various asynchronous, RDMA-based sparse times dense (SpMM) and sparse times sparse (SpGEMM) algorithms, evaluating their performance running in a distributed memory setting on GPUs. Our RDMA-based implementations use the NVSHMEM communication library for direct, asynchronous one-sided communication between GPUs. We compare our asynchronous implementations to state-of-the-art bulk synchronous GPU libraries as well as a CUDA-aware MPI implementation of the SUMMA algorithm. We find that asynchronous RDMA-based implementations are able to offer favorable performance compared to bulk synchronous implementations, while also allowing for the straightforward implementation of novel work stealing algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18141
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
Brock, Benjamin
Buluç, Aydın
Yelick, Katherine
Distributed, Parallel, and Cluster Computing
Sparse matrix multiplication is an important kernel for large-scale graph processing and other data-intensive applications. In this paper, we implement various asynchronous, RDMA-based sparse times dense (SpMM) and sparse times sparse (SpGEMM) algorithms, evaluating their performance running in a distributed memory setting on GPUs. Our RDMA-based implementations use the NVSHMEM communication library for direct, asynchronous one-sided communication between GPUs. We compare our asynchronous implementations to state-of-the-art bulk synchronous GPU libraries as well as a CUDA-aware MPI implementation of the SUMMA algorithm. We find that asynchronous RDMA-based implementations are able to offer favorable performance compared to bulk synchronous implementations, while also allowing for the straightforward implementation of novel work stealing algorithms.
title RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2311.18141