MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wan, Cheng, Tao, Runkai, Du, Zheng, Zhao, Yang Katie, Lin, Yingyan Celine
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915170988589056
author Wan, Cheng
Tao, Runkai
Du, Zheng
Zhao, Yang Katie
Lin, Yingyan Celine
author_facet Wan, Cheng
Tao, Runkai
Du, Zheng
Zhao, Yang Katie
Lin, Yingyan Celine
contents Graph convolutional networks (GCNs) have demonstrated superiority in graph-based learning tasks. However, training GCNs on full graphs is particularly challenging, due to the following two challenges: (1) the associated feature tensors can easily explode the memory and block the communication bandwidth of modern accelerators, and (2) the computation workflow in training GCNs alternates between sparse and dense matrix operations, complicating the efficient utilization of computational resources. Existing solutions for scalable distributed full-graph GCN training mostly adopt partition parallelism, which is unsatisfactory as they only partially address the first challenge while incurring scaled-out communication volume. To this end, we propose MixGCN aiming to simultaneously address both the aforementioned challenges towards GCN training. To tackle the first challenge, MixGCN integrates mixture of parallelism. Both theoretical and empirical analysis verify its constant communication volumes and enhanced balanced workload; For handling the second challenge, we consider mixture of accelerators (i.e., sparse and dense accelerators) with a dedicated accelerator for GCN training and a fine-grain pipeline. Extensive experiments show that MixGCN achieves boosted training efficiency and scalability.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01951
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators
Wan, Cheng
Tao, Runkai
Du, Zheng
Zhao, Yang Katie
Lin, Yingyan Celine
Machine Learning
Artificial Intelligence
Graph convolutional networks (GCNs) have demonstrated superiority in graph-based learning tasks. However, training GCNs on full graphs is particularly challenging, due to the following two challenges: (1) the associated feature tensors can easily explode the memory and block the communication bandwidth of modern accelerators, and (2) the computation workflow in training GCNs alternates between sparse and dense matrix operations, complicating the efficient utilization of computational resources. Existing solutions for scalable distributed full-graph GCN training mostly adopt partition parallelism, which is unsatisfactory as they only partially address the first challenge while incurring scaled-out communication volume. To this end, we propose MixGCN aiming to simultaneously address both the aforementioned challenges towards GCN training. To tackle the first challenge, MixGCN integrates mixture of parallelism. Both theoretical and empirical analysis verify its constant communication volumes and enhanced balanced workload; For handling the second challenge, we consider mixture of accelerators (i.e., sparse and dense accelerators) with a dedicated accelerator for GCN training and a fine-grain pipeline. Extensive experiments show that MixGCN achieves boosted training efficiency and scalability.
title MixGCN: Scalable GCN Training by Mixture of Parallelism and Mixture of Accelerators
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2501.01951