Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Qilong, Abdulah, Sameh, Abduljabbar, Mustafa, Ltaief, Hatem, Herten, Andreas, Bode, Mathis, Pratola, Matthew, Fadikar, Arindam, Genton, Marc G., Keyes, David E., Sun, Ying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914419361972224
author Pan, Qilong
Abdulah, Sameh
Abduljabbar, Mustafa
Ltaief, Hatem
Herten, Andreas
Bode, Mathis
Pratola, Matthew
Fadikar, Arindam
Genton, Marc G.
Keyes, David E.
Sun, Ying
author_facet Pan, Qilong
Abdulah, Sameh
Abduljabbar, Mustafa
Ltaief, Hatem
Herten, Andreas
Bode, Mathis
Pratola, Matthew
Fadikar, Arindam
Genton, Marc G.
Keyes, David E.
Sun, Ying
contents Emulating computationally intensive scientific simulations is crucial for enabling uncertainty quantification, optimization, and informed decision-making at scale. Gaussian Processes (GPs) offer a flexible and data-efficient foundation for statistical emulation, but their poor scalability limits applicability to large datasets. We introduce the Scaled Block Vecchia (SBV) algorithm for distributed GPU-based systems. SBV integrates the Scaled Vecchia approach for anisotropic input scaling with the Block Vecchia (BV) method to reduce computational and memory complexity while leveraging GPU acceleration techniques for efficient linear algebra operations. To the best of our knowledge, this is the first distributed implementation of any Vecchia-based GP variant. Our implementation employs MPI for inter-node parallelism and the MAGMA library for GPU-accelerated batched matrix computations. We demonstrate the scalability and efficiency of the proposed algorithm through experiments on synthetic and real-world workloads, including a 50M point simulation from a respiratory disease model. SBV achieves near-linear scalability on up to 512 A100 and GH200 GPUs, handles 2.56B points, and reduces energy use relative to exact GP solvers, establishing SBV as a scalable and energy-efficient framework for emulating large-scale scientific models on GPU-based distributed systems.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
Pan, Qilong
Abdulah, Sameh
Abduljabbar, Mustafa
Ltaief, Hatem
Herten, Andreas
Bode, Mathis
Pratola, Matthew
Fadikar, Arindam
Genton, Marc G.
Keyes, David E.
Sun, Ying
Distributed, Parallel, and Cluster Computing
Emulating computationally intensive scientific simulations is crucial for enabling uncertainty quantification, optimization, and informed decision-making at scale. Gaussian Processes (GPs) offer a flexible and data-efficient foundation for statistical emulation, but their poor scalability limits applicability to large datasets. We introduce the Scaled Block Vecchia (SBV) algorithm for distributed GPU-based systems. SBV integrates the Scaled Vecchia approach for anisotropic input scaling with the Block Vecchia (BV) method to reduce computational and memory complexity while leveraging GPU acceleration techniques for efficient linear algebra operations. To the best of our knowledge, this is the first distributed implementation of any Vecchia-based GP variant. Our implementation employs MPI for inter-node parallelism and the MAGMA library for GPU-accelerated batched matrix computations. We demonstrate the scalability and efficiency of the proposed algorithm through experiments on synthetic and real-world workloads, including a 50M point simulation from a respiratory disease model. SBV achieves near-linear scalability on up to 512 A100 and GH200 GPUs, handles 2.56B points, and reduces energy use relative to exact GP solvers, establishing SBV as a scalable and energy-efficient framework for emulating large-scale scientific models on GPU-based distributed systems.
title Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2504.12004