Rethinking Graph Transformer Architecture Design for Node Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Jiajun, Chen, Xuanze, Xie, Chenxuan, Shanqing, Yu, Xuan, Qi, Yang, Xiaoniu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916439752966144
author Zhou, Jiajun
Chen, Xuanze
Xie, Chenxuan
Shanqing, Yu
Xuan, Qi
Yang, Xiaoniu
author_facet Zhou, Jiajun
Chen, Xuanze
Xie, Chenxuan
Shanqing, Yu
Xuan, Qi
Yang, Xiaoniu
contents Graph Transformer (GT), as a special type of Graph Neural Networks (GNNs), utilizes multi-head attention to facilitate high-order message passing. However, this also imposes several limitations in node classification applications: 1) nodes are susceptible to global noise; 2) self-attention computation cannot scale well to large graphs. In this work, we conduct extensive observational experiments to explore the adaptability of the GT architecture in node classification tasks and draw several conclusions: the current multi-head self-attention module in GT can be completely replaceable, while the feed-forward neural network module proves to be valuable. Based on this, we decouple the propagation (P) and transformation (T) of GNNs and explore a powerful GT architecture, named GNNFormer, which is based on the P/T combination message passing and adapted for node classification in both homophilous and heterophilous scenarios. Extensive experiments on 12 benchmark datasets demonstrate that our proposed GT architecture can effectively adapt to node classification tasks without being affected by global noise and computational efficiency limitations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11189
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Graph Transformer Architecture Design for Node Classification
Zhou, Jiajun
Chen, Xuanze
Xie, Chenxuan
Shanqing, Yu
Xuan, Qi
Yang, Xiaoniu
Machine Learning
Graph Transformer (GT), as a special type of Graph Neural Networks (GNNs), utilizes multi-head attention to facilitate high-order message passing. However, this also imposes several limitations in node classification applications: 1) nodes are susceptible to global noise; 2) self-attention computation cannot scale well to large graphs. In this work, we conduct extensive observational experiments to explore the adaptability of the GT architecture in node classification tasks and draw several conclusions: the current multi-head self-attention module in GT can be completely replaceable, while the feed-forward neural network module proves to be valuable. Based on this, we decouple the propagation (P) and transformation (T) of GNNs and explore a powerful GT architecture, named GNNFormer, which is based on the P/T combination message passing and adapted for node classification in both homophilous and heterophilous scenarios. Extensive experiments on 12 benchmark datasets demonstrate that our proposed GT architecture can effectively adapt to node classification tasks without being affected by global noise and computational efficiency limitations.
title Rethinking Graph Transformer Architecture Design for Node Classification
topic Machine Learning
url https://arxiv.org/abs/2410.11189