HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Shujie, Wu, Yuxia, Shi, Chuan, Fang, Yuan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929745681186816
author Li, Shujie
Wu, Yuxia
Shi, Chuan
Fang, Yuan
author_facet Li, Shujie
Wu, Yuxia
Shi, Chuan
Fang, Yuan
contents Graph neural networks (GNNs) have demonstrated success in modeling relational data primarily under the assumption of homophily. However, many real-world graphs exhibit heterophily, where linked nodes belong to different categories or possess diverse attributes. Additionally, nodes in many domains are associated with textual descriptions, forming heterophilic text-attributed graphs (TAGs). Despite their significance, the study of heterophilic TAGs remains underexplored due to the lack of comprehensive benchmarks. To address this gap, we introduce the Heterophilic Text-attributed Graph Benchmark (HeTGB), a novel benchmark comprising five real-world heterophilic graph datasets from diverse domains, with nodes enriched by extensive textual descriptions. HeTGB enables systematic evaluation of GNNs, pre-trained language models (PLMs) and co-training methods on the node classification task. Through extensive benchmarking experiments, we showcase the utility of text attributes in heterophilic graphs, analyze the challenges posed by heterophilic TAGs and the limitations of existing models, and provide insights into the interplay between graph structures and textual attributes. We have publicly released HeTGB with baseline implementations to facilitate further research in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs
Li, Shujie
Wu, Yuxia
Shi, Chuan
Fang, Yuan
Computation and Language
Artificial Intelligence
Graph neural networks (GNNs) have demonstrated success in modeling relational data primarily under the assumption of homophily. However, many real-world graphs exhibit heterophily, where linked nodes belong to different categories or possess diverse attributes. Additionally, nodes in many domains are associated with textual descriptions, forming heterophilic text-attributed graphs (TAGs). Despite their significance, the study of heterophilic TAGs remains underexplored due to the lack of comprehensive benchmarks. To address this gap, we introduce the Heterophilic Text-attributed Graph Benchmark (HeTGB), a novel benchmark comprising five real-world heterophilic graph datasets from diverse domains, with nodes enriched by extensive textual descriptions. HeTGB enables systematic evaluation of GNNs, pre-trained language models (PLMs) and co-training methods on the node classification task. Through extensive benchmarking experiments, we showcase the utility of text attributes in heterophilic graphs, analyze the challenges posed by heterophilic TAGs and the limitations of existing models, and provide insights into the interplay between graph structures and textual attributes. We have publicly released HeTGB with baseline implementations to facilitate further research in this field.
title HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.04822