Communication-reduced Conjugate Gradient Variants for GPU-accelerated Clusters

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bernaschi, Massimo, Carrozzo, Mauro G., Celestini, Alessandro, Piperno, Giacomo, D'Ambra, Pasqua
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911594834821120
author Bernaschi, Massimo
Carrozzo, Mauro G.
Celestini, Alessandro
Piperno, Giacomo
D'Ambra, Pasqua
author_facet Bernaschi, Massimo
Carrozzo, Mauro G.
Celestini, Alessandro
Piperno, Giacomo
D'Ambra, Pasqua
contents Linear solvers are key components in any software platform for scientific and engineering computing. The solution of large and sparse linear systems lies at the core of physics-driven numerical simulations relying on partial differential equations (PDEs) and often represents a significant bottleneck in datadriven procedures, such as scientific machine learning. In this paper, we present an efficient implementation of the preconditioned s-step Conjugate Gradient (CG) method, originally proposed by Chronopoulos and Gear in 1989, for large clusters of Nvidia GPU-accelerated computing nodes. The method, often referred to as communication-reduced or communication-avoiding CG, reduces global synchronizations and data communication steps compared to the standard approach, enhancing strong and weak scalability on parallel computers. Our main contribution is the design of a parallel solver that fully exploits the aggregation of low-granularity operations inherent to the s-step CG method to leverage the high throughput of GPU accelerators. Additionally, it applies overlap between data communication and computation in the multi-GPU sparse matrix-vector product. Experiments on classic benchmark datasets, derived from the discretization of the Poisson PDE, demonstrate the potential of the method.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03743
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Communication-reduced Conjugate Gradient Variants for GPU-accelerated Clusters
Bernaschi, Massimo
Carrozzo, Mauro G.
Celestini, Alessandro
Piperno, Giacomo
D'Ambra, Pasqua
Numerical Analysis
65F10, 65Y05
G.4
Linear solvers are key components in any software platform for scientific and engineering computing. The solution of large and sparse linear systems lies at the core of physics-driven numerical simulations relying on partial differential equations (PDEs) and often represents a significant bottleneck in datadriven procedures, such as scientific machine learning. In this paper, we present an efficient implementation of the preconditioned s-step Conjugate Gradient (CG) method, originally proposed by Chronopoulos and Gear in 1989, for large clusters of Nvidia GPU-accelerated computing nodes. The method, often referred to as communication-reduced or communication-avoiding CG, reduces global synchronizations and data communication steps compared to the standard approach, enhancing strong and weak scalability on parallel computers. Our main contribution is the design of a parallel solver that fully exploits the aggregation of low-granularity operations inherent to the s-step CG method to leverage the high throughput of GPU accelerators. Additionally, it applies overlap between data communication and computation in the multi-GPU sparse matrix-vector product. Experiments on classic benchmark datasets, derived from the discretization of the Poisson PDE, demonstrate the potential of the method.
title Communication-reduced Conjugate Gradient Variants for GPU-accelerated Clusters
topic Numerical Analysis
65F10, 65Y05
G.4
url https://arxiv.org/abs/2501.03743