SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Xin, Zhuo, Weipeng, Vuong, Minh Phu, Li, Shiju, Kim, Jongryool, Rees, Bradley, Lee, Chul-Ho
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918234820706304
author Huang, Xin
Zhuo, Weipeng
Vuong, Minh Phu
Li, Shiju
Kim, Jongryool
Rees, Bradley
Lee, Chul-Ho
author_facet Huang, Xin
Zhuo, Weipeng
Vuong, Minh Phu
Li, Shiju
Kim, Jongryool
Rees, Bradley
Lee, Chul-Ho
contents Recently, distributed GNN training frameworks, such as DistDGL and PyG, have been developed to enable training GNN models on large graphs by leveraging multiple GPUs in a distributed manner. Despite these advances, their memory requirements are still excessively high, thereby hindering GNN training on large graphs using commodity workstations. In this paper, we propose SDT-GNN, a streaming-based distributed GNN training framework. Unlike the existing frameworks that load the entire graph in memory, it takes a stream of edges as input for graph partitioning to reduce the memory requirement for partitioning. It also enables distributed GNN training even when the aggregated memory size of GPUs is smaller than the size of the graph and feature data. Furthermore, to improve the quality of partitioning, we propose SPRING, a novel streaming partitioning algorithm for distributed GNN training. We demonstrate the effectiveness and efficiency of SDT-GNN on seven large public datasets. SDT-GNN has up to 95% less memory footprint than DistDGL and PyG without sacrificing the prediction accuracy. SPRING also outperforms state-of-the-art streaming partitioning algorithms significantly.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02300
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
Huang, Xin
Zhuo, Weipeng
Vuong, Minh Phu
Li, Shiju
Kim, Jongryool
Rees, Bradley
Lee, Chul-Ho
Machine Learning
Distributed, Parallel, and Cluster Computing
Recently, distributed GNN training frameworks, such as DistDGL and PyG, have been developed to enable training GNN models on large graphs by leveraging multiple GPUs in a distributed manner. Despite these advances, their memory requirements are still excessively high, thereby hindering GNN training on large graphs using commodity workstations. In this paper, we propose SDT-GNN, a streaming-based distributed GNN training framework. Unlike the existing frameworks that load the entire graph in memory, it takes a stream of edges as input for graph partitioning to reduce the memory requirement for partitioning. It also enables distributed GNN training even when the aggregated memory size of GPUs is smaller than the size of the graph and feature data. Furthermore, to improve the quality of partitioning, we propose SPRING, a novel streaming partitioning algorithm for distributed GNN training. We demonstrate the effectiveness and efficiency of SDT-GNN on seven large public datasets. SDT-GNN has up to 95% less memory footprint than DistDGL and PyG without sacrificing the prediction accuracy. SPRING also outperforms state-of-the-art streaming partitioning algorithms significantly.
title SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2404.02300