EmbedPart: Embedding-Driven Graph Partitioning for Scalable Graph Neural Network Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Merkel, Nikolai, Mayer, Ruben, Markl, Volker, Jacobsen, Hans-Arno
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910094323613696
author Merkel, Nikolai
Mayer, Ruben
Markl, Volker
Jacobsen, Hans-Arno
author_facet Merkel, Nikolai
Mayer, Ruben
Markl, Volker
Jacobsen, Hans-Arno
contents Graph Neural Networks (GNNs) are widely used for learning on graph-structured data, but scaling GNN training to massive graphs remains challenging. To enable scalable distributed training, graphs are divided into smaller partitions that are distributed across multiple machines such that inter-machine communication is minimized and computational load is balanced. In practice, existing partitioning approaches face a fundamental trade-off between partitioning overhead and partitioning quality. We propose EmbedPart, an embedding-driven partitioning approach that achieves both speed and quality. Instead of operating directly on irregular graph structures, EmbedPart leverages node embeddings produced during the actual GNN training workload and clusters these dense embeddings to derive a partitioning. EmbedPart achieves more than 100x speedup over Metis while maintaining competitive partitioning quality and accelerating distributed GNN training. Moreover, EmbedPart naturally supports graph updates and fast repartitioning, and can be applied to graph reordering to improve data locality and accelerate single-machine GNN training. By shifting partitioning from irregular graph structures to dense embeddings, EmbedPart enables scalable and high-quality graph data optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01000
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EmbedPart: Embedding-Driven Graph Partitioning for Scalable Graph Neural Network Training
Merkel, Nikolai
Mayer, Ruben
Markl, Volker
Jacobsen, Hans-Arno
Machine Learning
Databases
Distributed, Parallel, and Cluster Computing
Graph Neural Networks (GNNs) are widely used for learning on graph-structured data, but scaling GNN training to massive graphs remains challenging. To enable scalable distributed training, graphs are divided into smaller partitions that are distributed across multiple machines such that inter-machine communication is minimized and computational load is balanced. In practice, existing partitioning approaches face a fundamental trade-off between partitioning overhead and partitioning quality. We propose EmbedPart, an embedding-driven partitioning approach that achieves both speed and quality. Instead of operating directly on irregular graph structures, EmbedPart leverages node embeddings produced during the actual GNN training workload and clusters these dense embeddings to derive a partitioning. EmbedPart achieves more than 100x speedup over Metis while maintaining competitive partitioning quality and accelerating distributed GNN training. Moreover, EmbedPart naturally supports graph updates and fast repartitioning, and can be applied to graph reordering to improve data locality and accelerate single-machine GNN training. By shifting partitioning from irregular graph structures to dense embeddings, EmbedPart enables scalable and high-quality graph data optimization.
title EmbedPart: Embedding-Driven Graph Partitioning for Scalable Graph Neural Network Training
topic Machine Learning
Databases
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2604.01000