GPU-Resident Gaussian Process Regression Leveraging Asynchronous Tasks with HPX

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Möllmann, Henrik, Pflüger, Dirk, Strack, Alexander
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912920396365824
author Möllmann, Henrik
Pflüger, Dirk
Strack, Alexander
author_facet Möllmann, Henrik
Pflüger, Dirk
Strack, Alexander
contents Gaussian processes (GPs) are a widely used regression tool, but the cubic complexity of exact solvers limits their scalability. To address this challenge, we extend the GPRat library by incorporating a fully GPU-resident GP prediction pipeline. GPRat is an HPX-based library that combines task-based parallelism with an intuitive Python API. We implement tiled algorithms for the GP prediction using optimized CUDA libraries, thereby exploiting massive parallelism for linear algebra operations. We evaluate the optimal number of CUDA streams and compare the performance of our GPU implementation to the existing CPU-based implementation. Our results show the GPU implementation provides speedups for datasets larger than 128 training samples. We observe speedups of up to 4.3 for the Cholesky decomposition itself and 4.6 for the GP prediction. Furthermore, combining HPX with multiple CUDA streams allows GPRat to match, and for large datasets, surpass cuSOLVER's performance by up to 11 percent.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19683
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GPU-Resident Gaussian Process Regression Leveraging Asynchronous Tasks with HPX
Möllmann, Henrik
Pflüger, Dirk
Strack, Alexander
Distributed, Parallel, and Cluster Computing
Gaussian processes (GPs) are a widely used regression tool, but the cubic complexity of exact solvers limits their scalability. To address this challenge, we extend the GPRat library by incorporating a fully GPU-resident GP prediction pipeline. GPRat is an HPX-based library that combines task-based parallelism with an intuitive Python API. We implement tiled algorithms for the GP prediction using optimized CUDA libraries, thereby exploiting massive parallelism for linear algebra operations. We evaluate the optimal number of CUDA streams and compare the performance of our GPU implementation to the existing CPU-based implementation. Our results show the GPU implementation provides speedups for datasets larger than 128 training samples. We observe speedups of up to 4.3 for the Cholesky decomposition itself and 4.6 for the GP prediction. Furthermore, combining HPX with multiple CUDA streams allows GPRat to match, and for large datasets, surpass cuSOLVER's performance by up to 11 percent.
title GPU-Resident Gaussian Process Regression Leveraging Asynchronous Tasks with HPX
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.19683