RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916274097881088 |
|---|---|
| author | Brock, Benjamin Buluç, Aydın Yelick, Katherine |
| author_facet | Brock, Benjamin Buluç, Aydın Yelick, Katherine |
| contents | Sparse matrix multiplication is an important kernel for large-scale graph processing and other data-intensive applications. In this paper, we implement various asynchronous, RDMA-based sparse times dense (SpMM) and sparse times sparse (SpGEMM) algorithms, evaluating their performance running in a distributed memory setting on GPUs. Our RDMA-based implementations use the NVSHMEM communication library for direct, asynchronous one-sided communication between GPUs. We compare our asynchronous implementations to state-of-the-art bulk synchronous GPU libraries as well as a CUDA-aware MPI implementation of the SUMMA algorithm. We find that asynchronous RDMA-based implementations are able to offer favorable performance compared to bulk synchronous implementations, while also allowing for the straightforward implementation of novel work stealing algorithms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_18141 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs Brock, Benjamin Buluç, Aydın Yelick, Katherine Distributed, Parallel, and Cluster Computing Sparse matrix multiplication is an important kernel for large-scale graph processing and other data-intensive applications. In this paper, we implement various asynchronous, RDMA-based sparse times dense (SpMM) and sparse times sparse (SpGEMM) algorithms, evaluating their performance running in a distributed memory setting on GPUs. Our RDMA-based implementations use the NVSHMEM communication library for direct, asynchronous one-sided communication between GPUs. We compare our asynchronous implementations to state-of-the-art bulk synchronous GPU libraries as well as a CUDA-aware MPI implementation of the SUMMA algorithm. We find that asynchronous RDMA-based implementations are able to offer favorable performance compared to bulk synchronous implementations, while also allowing for the straightforward implementation of novel work stealing algorithms. |
| title | RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2311.18141 |