PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ockerman, Seth, Gueroudji, Amal, Mallick, Tanwi, He, Yixuan, Pouchard, Line, Ross, Robert, Venkataraman, Shivaram
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908540343418880
author Ockerman, Seth
Gueroudji, Amal
Mallick, Tanwi
He, Yixuan
Pouchard, Line
Ross, Robert
Venkataraman, Shivaram
author_facet Ockerman, Seth
Gueroudji, Amal
Mallick, Tanwi
He, Yixuan
Pouchard, Line
Ross, Robert
Venkataraman, Shivaram
contents Spatiotemporal graph neural networks (ST-GNNs) are powerful tools for modeling spatial and temporal data dependencies. However, their applications have been limited primarily to small-scale datasets because of memory constraints. While distributed training offers a solution, current frameworks lack support for spatiotemporal models and overlook the properties of spatiotemporal data. Informed by a scaling study on a large-scale workload, we present PyTorch Geometric Temporal Index (PGT-I), an extension to PyTorch Geometric Temporal that integrates distributed data parallel training and two novel strategies: index-batching and distributed-index-batching. Our index techniques exploit spatiotemporal structure to construct snapshots dynamically at runtime, significantly reducing memory overhead, while distributed-index-batching extends this approach by enabling scalable processing across multiple GPUs. Our techniques enable the first-ever training of an ST-GNN on the entire PeMS dataset without graph partitioning, reducing peak memory usage by up to 89% and achieving up to a 11.78x speedup over standard DDP with 128 GPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11683
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
Ockerman, Seth
Gueroudji, Amal
Mallick, Tanwi
He, Yixuan
Pouchard, Line
Ross, Robert
Venkataraman, Shivaram
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Spatiotemporal graph neural networks (ST-GNNs) are powerful tools for modeling spatial and temporal data dependencies. However, their applications have been limited primarily to small-scale datasets because of memory constraints. While distributed training offers a solution, current frameworks lack support for spatiotemporal models and overlook the properties of spatiotemporal data. Informed by a scaling study on a large-scale workload, we present PyTorch Geometric Temporal Index (PGT-I), an extension to PyTorch Geometric Temporal that integrates distributed data parallel training and two novel strategies: index-batching and distributed-index-batching. Our index techniques exploit spatiotemporal structure to construct snapshots dynamically at runtime, significantly reducing memory overhead, while distributed-index-batching extends this approach by enabling scalable processing across multiple GPUs. Our techniques enable the first-ever training of an ST-GNN on the entire PeMS dataset without graph partitioning, reducing peak memory usage by up to 89% and achieving up to a 11.78x speedup over standard DDP with 128 GPUs.
title PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.11683