Masked Language Models are Good Heterogeneous Graph Generalizers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jinyu, Yang, Cheng, Cui, Shanyuan, Guo, Zeyuan, Yang, Liangwei, Zhang, Muhan, Zhang, Zhiqiang, Shi, Chuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911083229347840
author Yang, Jinyu
Yang, Cheng
Cui, Shanyuan
Guo, Zeyuan
Yang, Liangwei
Zhang, Muhan
Zhang, Zhiqiang
Shi, Chuan
author_facet Yang, Jinyu
Yang, Cheng
Cui, Shanyuan
Guo, Zeyuan
Yang, Liangwei
Zhang, Muhan
Zhang, Zhiqiang
Shi, Chuan
contents Heterogeneous graph neural networks (HGNNs) excel at capturing structural and semantic information in heterogeneous graphs (HGs), while struggling to generalize across domains and tasks. With the rapid advancement of large language models (LLMs), a recent study explored the integration of HGNNs with LLMs for generalizable heterogeneous graph learning. However, this approach typically encodes structural information as HG tokens using HGNNs, and disparities in embedding spaces between HGNNs and LLMs have been shown to bias the LLM's comprehension of HGs. Moreover, since these HG tokens are often derived from node-level tasks, the model's ability to generalize across tasks remains limited. To this end, we propose a simple yet effective Masked Language Modeling-based method, called MLM4HG. MLM4HG introduces metapath-based textual sequences instead of HG tokens to extract structural and semantic information inherent in HGs, and designs customized textual templates to unify different graph tasks into a coherent cloze-style 'mask' token prediction paradigm. Specifically,MLM4HG first converts HGs from various domains to texts based on metapaths, and subsequently combines them with the unified task texts to form a HG-based corpus. Moreover, the corpus is fed into a pretrained LM for fine-tuning with a constrained target vocabulary, enabling the fine-tuned LM to generalize to unseen target HGs. Extensive cross-domain and multi-task experiments on four real-world datasets demonstrate the superior generalization performance of MLM4HG over state-of-the-art methods in both few-shot and zero-shot scenarios. Our code is available at https://github.com/BUPT-GAMMA/MLM4HG.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06157
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Masked Language Models are Good Heterogeneous Graph Generalizers
Yang, Jinyu
Yang, Cheng
Cui, Shanyuan
Guo, Zeyuan
Yang, Liangwei
Zhang, Muhan
Zhang, Zhiqiang
Shi, Chuan
Social and Information Networks
Computation and Language
Heterogeneous graph neural networks (HGNNs) excel at capturing structural and semantic information in heterogeneous graphs (HGs), while struggling to generalize across domains and tasks. With the rapid advancement of large language models (LLMs), a recent study explored the integration of HGNNs with LLMs for generalizable heterogeneous graph learning. However, this approach typically encodes structural information as HG tokens using HGNNs, and disparities in embedding spaces between HGNNs and LLMs have been shown to bias the LLM's comprehension of HGs. Moreover, since these HG tokens are often derived from node-level tasks, the model's ability to generalize across tasks remains limited. To this end, we propose a simple yet effective Masked Language Modeling-based method, called MLM4HG. MLM4HG introduces metapath-based textual sequences instead of HG tokens to extract structural and semantic information inherent in HGs, and designs customized textual templates to unify different graph tasks into a coherent cloze-style 'mask' token prediction paradigm. Specifically,MLM4HG first converts HGs from various domains to texts based on metapaths, and subsequently combines them with the unified task texts to form a HG-based corpus. Moreover, the corpus is fed into a pretrained LM for fine-tuning with a constrained target vocabulary, enabling the fine-tuned LM to generalize to unseen target HGs. Extensive cross-domain and multi-task experiments on four real-world datasets demonstrate the superior generalization performance of MLM4HG over state-of-the-art methods in both few-shot and zero-shot scenarios. Our code is available at https://github.com/BUPT-GAMMA/MLM4HG.
title Masked Language Models are Good Heterogeneous Graph Generalizers
topic Social and Information Networks
Computation and Language
url https://arxiv.org/abs/2506.06157