GraphAr: An Efficient Storage Scheme for Graph Data in Data Lakes

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Xue, Zeng, Weibin, Wang, Zhibin, Zhu, Diwen, Xu, Jingbo, Yu, Wenyuan, Zhou, Jingren
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916410295320576
author Li, Xue
Zeng, Weibin
Wang, Zhibin
Zhu, Diwen
Xu, Jingbo
Yu, Wenyuan
Zhou, Jingren
author_facet Li, Xue
Zeng, Weibin
Wang, Zhibin
Zhu, Diwen
Xu, Jingbo
Yu, Wenyuan
Zhou, Jingren
contents Data lakes, increasingly adopted for their ability to store and analyze diverse types of data, commonly use columnar storage formats like Parquet and ORC for handling relational tables. However, these traditional setups fall short when it comes to efficiently managing graph data, particularly those conforming to the Labeled Property Graph (LPG) model. To address this gap, this paper introduces GraphAr, a specialized storage scheme designed to enhance existing data lakes for efficient graph data management. Leveraging the strengths of Parquet, GraphAr captures LPG semantics precisely and facilitates graph-specific operations such as neighbor retrieval and label filtering. Through innovative data organization, encoding, and decoding techniques, GraphAr dramatically improves performance. Our evaluations reveal that GraphAr outperforms conventional Parquet and Acero-based methods, achieving an average speedup of 4452x for neighbor retrieval, 14.8x for label filtering, and 29.5x for end-to-end workloads. These findings highlight GraphAr's potential to extend the utility of data lakes by enabling efficient graph data management.
format Preprint
id arxiv_https___arxiv_org_abs_2312_09577
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle GraphAr: An Efficient Storage Scheme for Graph Data in Data Lakes
Li, Xue
Zeng, Weibin
Wang, Zhibin
Zhu, Diwen
Xu, Jingbo
Yu, Wenyuan
Zhou, Jingren
Databases
E.5; E.2; H.2.4; H.2.1
Data lakes, increasingly adopted for their ability to store and analyze diverse types of data, commonly use columnar storage formats like Parquet and ORC for handling relational tables. However, these traditional setups fall short when it comes to efficiently managing graph data, particularly those conforming to the Labeled Property Graph (LPG) model. To address this gap, this paper introduces GraphAr, a specialized storage scheme designed to enhance existing data lakes for efficient graph data management. Leveraging the strengths of Parquet, GraphAr captures LPG semantics precisely and facilitates graph-specific operations such as neighbor retrieval and label filtering. Through innovative data organization, encoding, and decoding techniques, GraphAr dramatically improves performance. Our evaluations reveal that GraphAr outperforms conventional Parquet and Acero-based methods, achieving an average speedup of 4452x for neighbor retrieval, 14.8x for label filtering, and 29.5x for end-to-end workloads. These findings highlight GraphAr's potential to extend the utility of data lakes by enabling efficient graph data management.
title GraphAr: An Efficient Storage Scheme for Graph Data in Data Lakes
topic Databases
E.5; E.2; H.2.4; H.2.1
url https://arxiv.org/abs/2312.09577