When Do Graph Foundation Models Transfer? A Data-Centric Theory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jiajun, Chen, Ying, Wang, Peihao, He, Yixuan, Li, Pan, Akella, Aditya, Wang, Zhangyang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916060984246272
author Zhu, Jiajun
Chen, Ying
Wang, Peihao
He, Yixuan
Li, Pan
Akella, Aditya
Wang, Zhangyang
author_facet Zhu, Jiajun
Chen, Ying
Wang, Peihao
He, Yixuan
Li, Pan
Akella, Aditya
Wang, Zhangyang
contents Graph foundation models (GFMs) aim to reuse a single backbone across diverse graph domains, yet their transfer is often uneven and can exhibit negative transfer. While most prior work improves transfer through architectural or adaptation choices, we ask a data-centric question: which properties of two graph domains determine how much a fixed representation model changes its outputs? Using a graphon-based continuous limit for dense graphs, we show that for both set-based and message-passing tokenizations, any Lipschitz backbone admits an explicit decomposition of cross-domain output shift into (i) graph-specific finite-sample approximation terms and (ii) an intrinsic, relabeling-invariant domain discrepancy capturing structural mismatch. A key ingredient is positional-encoding (PE) stability: we establish stability guarantees for spectral PEs and highlight contrasting behaviors of eigenvector- versus subspace-based PEs. Experiments on synthetic and real graphs validate the theory and translate the decomposition into guidance for data curation in GFM transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29828
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Do Graph Foundation Models Transfer? A Data-Centric Theory
Zhu, Jiajun
Chen, Ying
Wang, Peihao
He, Yixuan
Li, Pan
Akella, Aditya
Wang, Zhangyang
Machine Learning
Graph foundation models (GFMs) aim to reuse a single backbone across diverse graph domains, yet their transfer is often uneven and can exhibit negative transfer. While most prior work improves transfer through architectural or adaptation choices, we ask a data-centric question: which properties of two graph domains determine how much a fixed representation model changes its outputs? Using a graphon-based continuous limit for dense graphs, we show that for both set-based and message-passing tokenizations, any Lipschitz backbone admits an explicit decomposition of cross-domain output shift into (i) graph-specific finite-sample approximation terms and (ii) an intrinsic, relabeling-invariant domain discrepancy capturing structural mismatch. A key ingredient is positional-encoding (PE) stability: we establish stability guarantees for spectral PEs and highlight contrasting behaviors of eigenvector- versus subspace-based PEs. Experiments on synthetic and real graphs validate the theory and translate the decomposition into guidance for data curation in GFM transfer.
title When Do Graph Foundation Models Transfer? A Data-Centric Theory
topic Machine Learning
url https://arxiv.org/abs/2605.29828