Distributed Training of Large Graph Neural Networks with Variable Communication Rates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cervino, Juan, Turja, Md Asadullah, Mostafa, Hesham, Himayat, Nageen, Ribeiro, Alejandro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909231212396544
author Cervino, Juan
Turja, Md Asadullah
Mostafa, Hesham
Himayat, Nageen
Ribeiro, Alejandro
author_facet Cervino, Juan
Turja, Md Asadullah
Mostafa, Hesham
Himayat, Nageen
Ribeiro, Alejandro
contents Training Graph Neural Networks (GNNs) on large graphs presents unique challenges due to the large memory and computing requirements. Distributed GNN training, where the graph is partitioned across multiple machines, is a common approach to training GNNs on large graphs. However, as the graph cannot generally be decomposed into small non-interacting components, data communication between the training machines quickly limits training speeds. Compressing the communicated node activations by a fixed amount improves the training speeds, but lowers the accuracy of the trained GNN. In this paper, we introduce a variable compression scheme for reducing the communication volume in distributed GNN training without compromising the accuracy of the learned model. Based on our theoretical analysis, we derive a variable compression method that converges to a solution equivalent to the full communication case, for all graph partitioning schemes. Our empirical results show that our method attains a comparable performance to the one obtained with full communication. We outperform full communication at any fixed compression ratio for any communication budget.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17611
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Distributed Training of Large Graph Neural Networks with Variable Communication Rates
Cervino, Juan
Turja, Md Asadullah
Mostafa, Hesham
Himayat, Nageen
Ribeiro, Alejandro
Machine Learning
Signal Processing
Training Graph Neural Networks (GNNs) on large graphs presents unique challenges due to the large memory and computing requirements. Distributed GNN training, where the graph is partitioned across multiple machines, is a common approach to training GNNs on large graphs. However, as the graph cannot generally be decomposed into small non-interacting components, data communication between the training machines quickly limits training speeds. Compressing the communicated node activations by a fixed amount improves the training speeds, but lowers the accuracy of the trained GNN. In this paper, we introduce a variable compression scheme for reducing the communication volume in distributed GNN training without compromising the accuracy of the learned model. Based on our theoretical analysis, we derive a variable compression method that converges to a solution equivalent to the full communication case, for all graph partitioning schemes. Our empirical results show that our method attains a comparable performance to the one obtained with full communication. We outperform full communication at any fixed compression ratio for any communication budget.
title Distributed Training of Large Graph Neural Networks with Variable Communication Rates
topic Machine Learning
Signal Processing
url https://arxiv.org/abs/2406.17611