TGB 2.0: A Benchmark for Learning on Temporal Knowledge Graphs and Heterogeneous Graphs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gastinger, Julia, Huang, Shenyang, Galkin, Mikhail, Loghmani, Erfan, Parviz, Ali, Poursafaei, Farimah, Danovitch, Jacob, Rossi, Emanuele, Koutis, Ioannis, Stuckenschmidt, Heiner, Rabbany, Reihaneh, Rabusseau, Guillaume
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929548840402944
author Gastinger, Julia
Huang, Shenyang
Galkin, Mikhail
Loghmani, Erfan
Parviz, Ali
Poursafaei, Farimah
Danovitch, Jacob
Rossi, Emanuele
Koutis, Ioannis
Stuckenschmidt, Heiner
Rabbany, Reihaneh
Rabusseau, Guillaume
author_facet Gastinger, Julia
Huang, Shenyang
Galkin, Mikhail
Loghmani, Erfan
Parviz, Ali
Poursafaei, Farimah
Danovitch, Jacob
Rossi, Emanuele
Koutis, Ioannis
Stuckenschmidt, Heiner
Rabbany, Reihaneh
Rabusseau, Guillaume
contents Multi-relational temporal graphs are powerful tools for modeling real-world data, capturing the evolving and interconnected nature of entities over time. Recently, many novel models are proposed for ML on such graphs intensifying the need for robust evaluation and standardized benchmark datasets. However, the availability of such resources remains scarce and evaluation faces added complexity due to reproducibility issues in experimental protocols. To address these challenges, we introduce Temporal Graph Benchmark 2.0 (TGB 2.0), a novel benchmarking framework tailored for evaluating methods for predicting future links on Temporal Knowledge Graphs and Temporal Heterogeneous Graphs with a focus on large-scale datasets, extending the Temporal Graph Benchmark. TGB 2.0 facilitates comprehensive evaluations by presenting eight novel datasets spanning five domains with up to 53 million edges. TGB 2.0 datasets are significantly larger than existing datasets in terms of number of nodes, edges, or timestamps. In addition, TGB 2.0 provides a reproducible and realistic evaluation pipeline for multi-relational temporal graphs. Through extensive experimentation, we observe that 1) leveraging edge-type information is crucial to obtain high performance, 2) simple heuristic baselines are often competitive with more complex methods, 3) most methods fail to run on our largest datasets, highlighting the need for research on more scalable methods.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09639
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TGB 2.0: A Benchmark for Learning on Temporal Knowledge Graphs and Heterogeneous Graphs
Gastinger, Julia
Huang, Shenyang
Galkin, Mikhail
Loghmani, Erfan
Parviz, Ali
Poursafaei, Farimah
Danovitch, Jacob
Rossi, Emanuele
Koutis, Ioannis
Stuckenschmidt, Heiner
Rabbany, Reihaneh
Rabusseau, Guillaume
Machine Learning
Social and Information Networks
Multi-relational temporal graphs are powerful tools for modeling real-world data, capturing the evolving and interconnected nature of entities over time. Recently, many novel models are proposed for ML on such graphs intensifying the need for robust evaluation and standardized benchmark datasets. However, the availability of such resources remains scarce and evaluation faces added complexity due to reproducibility issues in experimental protocols. To address these challenges, we introduce Temporal Graph Benchmark 2.0 (TGB 2.0), a novel benchmarking framework tailored for evaluating methods for predicting future links on Temporal Knowledge Graphs and Temporal Heterogeneous Graphs with a focus on large-scale datasets, extending the Temporal Graph Benchmark. TGB 2.0 facilitates comprehensive evaluations by presenting eight novel datasets spanning five domains with up to 53 million edges. TGB 2.0 datasets are significantly larger than existing datasets in terms of number of nodes, edges, or timestamps. In addition, TGB 2.0 provides a reproducible and realistic evaluation pipeline for multi-relational temporal graphs. Through extensive experimentation, we observe that 1) leveraging edge-type information is crucial to obtain high performance, 2) simple heuristic baselines are often competitive with more complex methods, 3) most methods fail to run on our largest datasets, highlighting the need for research on more scalable methods.
title TGB 2.0: A Benchmark for Learning on Temporal Knowledge Graphs and Heterogeneous Graphs
topic Machine Learning
Social and Information Networks
url https://arxiv.org/abs/2406.09639