Multi-Scale Heterogeneous Text-Attributed Graph Datasets From Diverse Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yunhui, Xie, Qizhuo, Shi, Jinwei, Shen, Jiaxu, He, Tieke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916519824326656
author Liu, Yunhui
Xie, Qizhuo
Shi, Jinwei
Shen, Jiaxu
He, Tieke
author_facet Liu, Yunhui
Xie, Qizhuo
Shi, Jinwei
Shen, Jiaxu
He, Tieke
contents Heterogeneous Text-Attributed Graphs (HTAGs), where different types of entities are not only associated with texts but also connected by diverse relationships, have gained widespread popularity and application across various domains. However, current research on text-attributed graph learning predominantly focuses on homogeneous graphs, which feature a single node and edge type, thus leaving a gap in understanding how methods perform on HTAGs. One crucial reason is the lack of comprehensive HTAG datasets that offer original textual content and span multiple domains of varying sizes. To this end, we introduce a collection of challenging and diverse benchmark datasets for realistic and reproducible evaluation of machine learning models on HTAGs. Our HTAG datasets are multi-scale, span years in duration, and cover a wide range of domains, including movie, community question answering, academic, literature, and patent networks. We further conduct benchmark experiments on these datasets with various graph neural networks. All source data, dataset construction codes, processed HTAGs, data loaders, benchmark codes, and evaluation setup are publicly available at GitHub and Hugging Face.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08937
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Scale Heterogeneous Text-Attributed Graph Datasets From Diverse Domains
Liu, Yunhui
Xie, Qizhuo
Shi, Jinwei
Shen, Jiaxu
He, Tieke
Machine Learning
Computation and Language
Heterogeneous Text-Attributed Graphs (HTAGs), where different types of entities are not only associated with texts but also connected by diverse relationships, have gained widespread popularity and application across various domains. However, current research on text-attributed graph learning predominantly focuses on homogeneous graphs, which feature a single node and edge type, thus leaving a gap in understanding how methods perform on HTAGs. One crucial reason is the lack of comprehensive HTAG datasets that offer original textual content and span multiple domains of varying sizes. To this end, we introduce a collection of challenging and diverse benchmark datasets for realistic and reproducible evaluation of machine learning models on HTAGs. Our HTAG datasets are multi-scale, span years in duration, and cover a wide range of domains, including movie, community question answering, academic, literature, and patent networks. We further conduct benchmark experiments on these datasets with various graph neural networks. All source data, dataset construction codes, processed HTAGs, data loaders, benchmark codes, and evaluation setup are publicly available at GitHub and Hugging Face.
title Multi-Scale Heterogeneous Text-Attributed Graph Datasets From Diverse Domains
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2412.08937