GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Jaeyong, Park, Seongyeon, Jang, Hongsun, Jung, Jaewon, Lim, Hunseong, Hong, Junguk, Lee, Jinho
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918496316686336
author Song, Jaeyong
Park, Seongyeon
Jang, Hongsun
Jung, Jaewon
Lim, Hunseong
Hong, Junguk
Lee, Jinho
author_facet Song, Jaeyong
Park, Seongyeon
Jang, Hongsun
Jung, Jaewon
Lim, Hunseong
Hong, Junguk
Lee, Jinho
contents Full-graph training of graph neural networks (GNNs) is widely used as it enables direct validation of algorithmic improvements by preserving complete neighborhood information. However, it typically requires multiple GPUs or servers, incurring substantial hardware and inter-device communication costs. While existing single-server methods reduce infrastructure requirements, they remain constrained by GPU and host memory capacity as graph sizes increase. To address this limitation, we introduce GriNNder, which is the first work to leverage storage devices to enable full-graph training even with limited memory. Because modern NVMe SSDs offer multi-terabyte capacities and bandwidths exceeding 10 GB/s, they provide an appealing option when memory resources are scarce. Yet, directly applying storage-based methods from other domains fails to address the unique access patterns and data dependencies in full-graph GNN training. GriNNder tackles these challenges by structured storage offloading (SSO), a framework that manages the GPU-host-storage hierarchy through coordinated cache, (re)gather, and bypass mechanisms. To realize the framework, we devise (i) a partition-wise caching strategy for host memory that exploits the observation on cross-partition dependencies, (ii) a regathering strategy for gradient computation that eliminates redundant storage operations, and (iii) a lightweight partitioning scheme that mitigates the memory requirements of existing graph partitioners. In experiments performed over various models and datasets, GriNNder achieves up to 9.78x speedup over state-of-the-art baselines and throughput comparable to distributed systems, enabling previously infeasible large-scale full-graph training even on a single GPU.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11517
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
Song, Jaeyong
Park, Seongyeon
Jang, Hongsun
Jung, Jaewon
Lim, Hunseong
Hong, Junguk
Lee, Jinho
Distributed, Parallel, and Cluster Computing
Full-graph training of graph neural networks (GNNs) is widely used as it enables direct validation of algorithmic improvements by preserving complete neighborhood information. However, it typically requires multiple GPUs or servers, incurring substantial hardware and inter-device communication costs. While existing single-server methods reduce infrastructure requirements, they remain constrained by GPU and host memory capacity as graph sizes increase. To address this limitation, we introduce GriNNder, which is the first work to leverage storage devices to enable full-graph training even with limited memory. Because modern NVMe SSDs offer multi-terabyte capacities and bandwidths exceeding 10 GB/s, they provide an appealing option when memory resources are scarce. Yet, directly applying storage-based methods from other domains fails to address the unique access patterns and data dependencies in full-graph GNN training. GriNNder tackles these challenges by structured storage offloading (SSO), a framework that manages the GPU-host-storage hierarchy through coordinated cache, (re)gather, and bypass mechanisms. To realize the framework, we devise (i) a partition-wise caching strategy for host memory that exploits the observation on cross-partition dependencies, (ii) a regathering strategy for gradient computation that eliminates redundant storage operations, and (iii) a lightweight partitioning scheme that mitigates the memory requirements of existing graph partitioners. In experiments performed over various models and datasets, GriNNder achieves up to 9.78x speedup over state-of-the-art baselines and throughput comparable to distributed systems, enabling previously infeasible large-scale full-graph training even on a single GPU.
title GriNNder: Breaking the Memory Capacity Wall in Full-Graph GNN Training with Storage Offloading
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2605.11517