Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Qilong, Abdulah, Sameh, Abduljabbar, Mustafa, Ltaief, Hatem, Herten, Andreas, Bode, Mathis, Pratola, Matthew, Fadikar, Arindam, Genton, Marc G., Keyes, David E., Sun, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPU-Accelerated Vecchia Approximations of Gaussian Processes for Geospatial Data using Batched Matrix Computations
by: Pan, Qilong, et al.
Published: (2024)
by: Pan, Qilong, et al.
Published: (2024)
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
by: Geng, Zipei, et al.
Published: (2025)
by: Geng, Zipei, et al.
Published: (2025)
Parallel Approximations for High-Dimensional Multivariate Normal Probability Computation in Confidence Region Detection Applications
by: Zhang, Xiran, et al.
Published: (2024)
by: Zhang, Xiran, et al.
Published: (2024)
Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
High-Performance Statistical Computing (HPSC): Challenges, Opportunities, and Future Directions
by: Abdulah, Sameh, et al.
Published: (2025)
by: Abdulah, Sameh, et al.
Published: (2025)
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
by: Ahmed, Mahmoud, et al.
Published: (2026)
by: Ahmed, Mahmoud, et al.
Published: (2026)
Block Vecchia Approximation for Scalable and Efficient Gaussian Process Computations
by: Pan, Qilong, et al.
Published: (2024)
by: Pan, Qilong, et al.
Published: (2024)
exaCB: Reproducible Continuous Benchmark Collections at Scale Leveraging an Incremental Approach
by: Badwaik, Jayesh, et al.
Published: (2026)
by: Badwaik, Jayesh, et al.
Published: (2026)
RCOMPSs: A Scalable Runtime System for R Code Execution on Manycore Systems
by: Zhang, Xiran, et al.
Published: (2025)
by: Zhang, Xiran, et al.
Published: (2025)
A Novel Approach to Translate Structural Aggregation Queries to MapReduce Code
by: Abdelmoniem, Ahmed M., et al.
Published: (2025)
by: Abdelmoniem, Ahmed M., et al.
Published: (2025)
Efficient Accelerated Graph Edit Distance Computation on GPU
by: Dabah, Adel, et al.
Published: (2026)
by: Dabah, Adel, et al.
Published: (2026)
FedPBS: Proximal-Balanced Scaling Federated Learning Model for Robust Personalized Training for Non-IID Data
by: AbouNassar, Eman M., et al.
Published: (2026)
by: AbouNassar, Eman M., et al.
Published: (2026)
BlockEmulator: An Emulator Enabling to Test Blockchain Sharding Protocols
by: Huang, Huawei, et al.
Published: (2023)
by: Huang, Huawei, et al.
Published: (2023)
Round and Resilience-Optimal Approximate Agreement on Trees and Block Graphs
by: Fuchs, Marc, et al.
Published: (2025)
by: Fuchs, Marc, et al.
Published: (2025)
Universal Quantum Computer Simulation of 50 Qubits on Europe`s First Exascale Supercomputer Harnessing Its Heterogeneous CPU-GPU Architecture
by: De Raedt, Hans, et al.
Published: (2025)
by: De Raedt, Hans, et al.
Published: (2025)
Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search
by: Li, Zhonggen, et al.
Published: (2025)
by: Li, Zhonggen, et al.
Published: (2025)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
by: Cui, Shengkun, et al.
Published: (2025)
by: Cui, Shengkun, et al.
Published: (2025)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
by: Xi, Shaoke, et al.
Published: (2026)
by: Xi, Shaoke, et al.
Published: (2026)
CB-SpMV:A Data Aggregating and Balance Algorithm for Cache-Friendly Block-Based SpMV on GPUs
by: Cong, Xing, et al.
Published: (2026)
by: Cong, Xing, et al.
Published: (2026)
DeepOps & SLURM: Your GPU Cluster Guide
by: Majee, Arindam
Published: (2024)
by: Majee, Arindam
Published: (2024)
Sketched Gaussian Mechanism for Private Federated Learning
by: Li, Qiaobo, et al.
Published: (2025)
by: Li, Qiaobo, et al.
Published: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024)
by: Hsieh, Chou-Ying, et al.
Published: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
by: Almasri, Mohammad, et al.
Published: (2022)
by: Almasri, Mohammad, et al.
Published: (2022)
Optimizing sDTW for AMD GPUs
by: Latta-Lin, Daniel, et al.
Published: (2024)
by: Latta-Lin, Daniel, et al.
Published: (2024)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Serving Compound Inference Systems on Datacenter GPUs
by: Devata, Sriram, et al.
Published: (2026)
by: Devata, Sriram, et al.
Published: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
by: Jangda, Abhinav, et al.
Published: (2024)
by: Jangda, Abhinav, et al.
Published: (2024)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
Tensor-Parallel Emulation of Quantum Circuits with Block-Cyclic Distributed Matrix Product States
by: Adamski, Jakub, et al.
Published: (2025)
by: Adamski, Jakub, et al.
Published: (2025)
High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
by: Shi, Ruimin, et al.
Published: (2026)
by: Shi, Ruimin, et al.
Published: (2026)
Demystifying the Communication Characteristics for Distributed Transformer Models
by: Anthony, Quentin, et al.
Published: (2024)
by: Anthony, Quentin, et al.
Published: (2024)
Why Smaller Is Slower? Dimensional Misalignment in Compressed LLMs
by: Xin, Jihao, et al.
Published: (2026)
by: Xin, Jihao, et al.
Published: (2026)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
by: Plesner, Andreas, et al.
Published: (2024)
by: Plesner, Andreas, et al.
Published: (2024)
Managing Multi Instance GPUs for High Throughput and Energy Savings
by: Saraha, Abhijeet, et al.
Published: (2025)
by: Saraha, Abhijeet, et al.
Published: (2025)
Analytical Performance Estimation during Code Generation on Modern GPUs
by: Ernst, Dominik, et al.
Published: (2022)
by: Ernst, Dominik, et al.
Published: (2022)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
How to Rent GPUs on a Budget
by: Li, Zhouzi, et al.
Published: (2024)
by: Li, Zhouzi, et al.
Published: (2024)
Similar Items
-
GPU-Accelerated Vecchia Approximations of Gaussian Processes for Geospatial Data using Batched Matrix Computations
by: Pan, Qilong, et al.
Published: (2024) -
GPU-Accelerated Modified Bessel Function of the Second Kind for Gaussian Processes
by: Geng, Zipei, et al.
Published: (2025) -
Parallel Approximations for High-Dimensional Multivariate Normal Probability Computation in Confidence Region Detection Applications
by: Zhang, Xiran, et al.
Published: (2024) -
Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling
by: Ren, Jie, et al.
Published: (2024) -
High-Performance Statistical Computing (HPSC): Challenges, Opportunities, and Future Directions
by: Abdulah, Sameh, et al.
Published: (2025)