Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jing, Zhou, Yuhang, Qian, Shengyi, He, Zhongmou, Zhao, Tong, Shah, Neil, Koutra, Danai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910898601328640
author Zhu, Jing
Zhou, Yuhang
Qian, Shengyi
He, Zhongmou
Zhao, Tong
Shah, Neil
Koutra, Danai
author_facet Zhu, Jing
Zhou, Yuhang
Qian, Shengyi
He, Zhongmou
Zhao, Tong
Shah, Neil
Koutra, Danai
contents Graph machine learning has made significant strides in recent years, yet the integration of visual information with graph structure and its potential for improving performance in downstream tasks remains an underexplored area. To address this critical gap, we introduce the Multimodal Graph Benchmark (MM-GRAPH), a pioneering benchmark that incorporates both visual and textual information into graph learning tasks. MM-GRAPH extends beyond existing text-attributed graph benchmarks, offering a more comprehensive evaluation framework for multimodal graph learning Our benchmark comprises seven diverse datasets of varying scales (ranging from thousands to millions of edges), designed to assess algorithms across different tasks in real-world scenarios. These datasets feature rich multimodal node attributes, including visual data, which enables a more holistic evaluation of various graph learning frameworks in complex, multimodal environments. To support advancements in this emerging field, we provide an extensive empirical study on various graph learning frameworks when presented with features from multiple modalities, particularly emphasizing the impact of visual information. This study offers valuable insights into the challenges and opportunities of integrating visual data into graph learning.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning
Zhu, Jing
Zhou, Yuhang
Qian, Shengyi
He, Zhongmou
Zhao, Tong
Shah, Neil
Koutra, Danai
Machine Learning
Artificial Intelligence
Graph machine learning has made significant strides in recent years, yet the integration of visual information with graph structure and its potential for improving performance in downstream tasks remains an underexplored area. To address this critical gap, we introduce the Multimodal Graph Benchmark (MM-GRAPH), a pioneering benchmark that incorporates both visual and textual information into graph learning tasks. MM-GRAPH extends beyond existing text-attributed graph benchmarks, offering a more comprehensive evaluation framework for multimodal graph learning Our benchmark comprises seven diverse datasets of varying scales (ranging from thousands to millions of edges), designed to assess algorithms across different tasks in real-world scenarios. These datasets feature rich multimodal node attributes, including visual data, which enables a more holistic evaluation of various graph learning frameworks in complex, multimodal environments. To support advancements in this emerging field, we provide an extensive empirical study on various graph learning frameworks when presented with features from multiple modalities, particularly emphasizing the impact of visual information. This study offers valuable insights into the challenges and opportunities of integrating visual data into graph learning.
title Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.16321