PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Jigao, Boeschen, Nils, El-Hindi, Muhammad, Binnig, Carsten
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917513043902464
author Luo, Jigao
Boeschen, Nils
El-Hindi, Muhammad
Binnig, Carsten
author_facet Luo, Jigao
Boeschen, Nils
El-Hindi, Muhammad
Binnig, Carsten
contents The AI hardware boom has led modern data centers to adopt HPC-style architectures centered on distributed, GPU-centric computation. Large GPU clusters interconnected by fast RDMA networks and backed by high-bandwidth NVMe storage enable scalable computation and rapid access to storage-resident data. Tensor computation runtimes (TCRs), such as PyTorch, originally designed for AI workloads, have recently been shown to accelerate analytical workloads. However, prior work has primarily considered settings where the data fits in aggregated GPU memory. In this paper, we systematically study how TCRs can support scalable, distributed query processing for large-scale, storage-resident OLAP workloads. Although TCRs provide abstractions for network and storage I/O, naive use often underutilizes GPU and I/O bandwidth due to insufficient overlap between computation and data movement. As a core contribution, we present PystachIO, a prototype of a PyTorch-based distributed OLAP engine that combines fast network and storage I/O with key optimizations to maximize GPU, network, and storage utilization. Our evaluation shows up to 3x end-to-end speedups over existing distributed GPU-based query processing approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
Luo, Jigao
Boeschen, Nils
El-Hindi, Muhammad
Binnig, Carsten
Databases
The AI hardware boom has led modern data centers to adopt HPC-style architectures centered on distributed, GPU-centric computation. Large GPU clusters interconnected by fast RDMA networks and backed by high-bandwidth NVMe storage enable scalable computation and rapid access to storage-resident data. Tensor computation runtimes (TCRs), such as PyTorch, originally designed for AI workloads, have recently been shown to accelerate analytical workloads. However, prior work has primarily considered settings where the data fits in aggregated GPU memory. In this paper, we systematically study how TCRs can support scalable, distributed query processing for large-scale, storage-resident OLAP workloads. Although TCRs provide abstractions for network and storage I/O, naive use often underutilizes GPU and I/O bandwidth due to insufficient overlap between computation and data movement. As a core contribution, we present PystachIO, a prototype of a PyTorch-based distributed OLAP engine that combines fast network and storage I/O with key optimizations to maximize GPU, network, and storage utilization. Our evaluation shows up to 3x end-to-end speedups over existing distributed GPU-based query processing approaches.
title PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
topic Databases
url https://arxiv.org/abs/2512.02862