Saved in:
Bibliographic Details
Main Authors: Xu, Chongyang, Siebenbrunner, Christoph, Bindschaedler, Laurent
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.01872
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918342040748032
author Xu, Chongyang
Siebenbrunner, Christoph
Bindschaedler, Laurent
author_facet Xu, Chongyang
Siebenbrunner, Christoph
Bindschaedler, Laurent
contents Cross-partition edges dominate the cost of distributed GNN training: fetching remote features and activations per iteration overwhelms the network as graphs deepen and partition counts grow. Grappa is a distributed GNN training framework that enforces gradient-only communication: during each iteration, partitions train in isolation and exchange only gradients for the global update. To recover accuracy lost to isolation, Grappa (i) periodically repartitions to expose new neighborhoods and (ii) applies a lightweight coverage-corrected gradient aggregation inspired by importance sampling. We present an asymptotically unbiased estimator for gradient correction, which we use to develop a minimum-distance batch-level variant that is compatible with common deep-learning packages. We also introduce a shrinkage version that improves stability in practice. Empirical results on real and synthetic graphs show that Grappa trains GNNs 4x faster on average (up to 13x) than state-of-the-art systems, achieves better accuracy especially for deeper models, and sustains training at the trillion-edge scale on commodity hardware. Grappa is model-agnostic, supports full-graph and mini-batch training, and does not rely on high-bandwidth interconnects or caching.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01872
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Grappa: Gradient-Only Communication for Scalable Graph Neural Network Training
Xu, Chongyang
Siebenbrunner, Christoph
Bindschaedler, Laurent
Distributed, Parallel, and Cluster Computing
Machine Learning
Cross-partition edges dominate the cost of distributed GNN training: fetching remote features and activations per iteration overwhelms the network as graphs deepen and partition counts grow. Grappa is a distributed GNN training framework that enforces gradient-only communication: during each iteration, partitions train in isolation and exchange only gradients for the global update. To recover accuracy lost to isolation, Grappa (i) periodically repartitions to expose new neighborhoods and (ii) applies a lightweight coverage-corrected gradient aggregation inspired by importance sampling. We present an asymptotically unbiased estimator for gradient correction, which we use to develop a minimum-distance batch-level variant that is compatible with common deep-learning packages. We also introduce a shrinkage version that improves stability in practice. Empirical results on real and synthetic graphs show that Grappa trains GNNs 4x faster on average (up to 13x) than state-of-the-art systems, achieves better accuracy especially for deeper models, and sustains training at the trillion-edge scale on commodity hardware. Grappa is model-agnostic, supports full-graph and mini-batch training, and does not rely on high-bandwidth interconnects or caching.
title Grappa: Gradient-Only Communication for Scalable Graph Neural Network Training
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2602.01872