Virtual Nodes Can Help: Tackling Distribution Shifts in Federated Graph Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fu, Xingbo, Chen, Zihan, He, Yinhan, Wang, Song, Zhang, Binchi, Chen, Chen, Li, Jundong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929723471298560
author Fu, Xingbo
Chen, Zihan
He, Yinhan
Wang, Song
Zhang, Binchi
Chen, Chen
Li, Jundong
author_facet Fu, Xingbo
Chen, Zihan
He, Yinhan
Wang, Song
Zhang, Binchi
Chen, Chen
Li, Jundong
contents Federated Graph Learning (FGL) enables multiple clients to jointly train powerful graph learning models, e.g., Graph Neural Networks (GNNs), without sharing their local graph data for graph-related downstream tasks, such as graph property prediction. In the real world, however, the graph data can suffer from significant distribution shifts across clients as the clients may collect their graph data for different purposes. In particular, graph properties are usually associated with invariant label-relevant substructures (i.e., subgraphs) across clients, while label-irrelevant substructures can appear in a client-specific manner. The issue of distribution shifts of graph data hinders the efficiency of GNN training and leads to serious performance degradation in FGL. To tackle the aforementioned issue, we propose a novel FGL framework entitled FedVN that eliminates distribution shifts through client-specific graph augmentation strategies with multiple learnable Virtual Nodes (VNs). Specifically, FedVN lets the clients jointly learn a set of shared VNs while training a global GNN model. To eliminate distribution shifts, each client trains a personalized edge generator that determines how the VNs connect local graphs in a client-specific manner. Furthermore, we provide theoretical analyses indicating that FedVN can eliminate distribution shifts of graph data across clients. Comprehensive experiments on four datasets under five settings demonstrate the superiority of our proposed FedVN over nine baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19229
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Virtual Nodes Can Help: Tackling Distribution Shifts in Federated Graph Learning
Fu, Xingbo
Chen, Zihan
He, Yinhan
Wang, Song
Zhang, Binchi
Chen, Chen
Li, Jundong
Machine Learning
Distributed, Parallel, and Cluster Computing
Federated Graph Learning (FGL) enables multiple clients to jointly train powerful graph learning models, e.g., Graph Neural Networks (GNNs), without sharing their local graph data for graph-related downstream tasks, such as graph property prediction. In the real world, however, the graph data can suffer from significant distribution shifts across clients as the clients may collect their graph data for different purposes. In particular, graph properties are usually associated with invariant label-relevant substructures (i.e., subgraphs) across clients, while label-irrelevant substructures can appear in a client-specific manner. The issue of distribution shifts of graph data hinders the efficiency of GNN training and leads to serious performance degradation in FGL. To tackle the aforementioned issue, we propose a novel FGL framework entitled FedVN that eliminates distribution shifts through client-specific graph augmentation strategies with multiple learnable Virtual Nodes (VNs). Specifically, FedVN lets the clients jointly learn a set of shared VNs while training a global GNN model. To eliminate distribution shifts, each client trains a personalized edge generator that determines how the VNs connect local graphs in a client-specific manner. Furthermore, we provide theoretical analyses indicating that FedVN can eliminate distribution shifts of graph data across clients. Comprehensive experiments on four datasets under five settings demonstrate the superiority of our proposed FedVN over nine baselines.
title Virtual Nodes Can Help: Tackling Distribution Shifts in Federated Graph Learning
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2412.19229