Data Formats in Analytical DBMSs: Performance Trade-offs and Future Directions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Chunwei, Pavlenko, Anna, Interlandi, Matteo, Haynes, Brandon
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909398421471232
author Liu, Chunwei
Pavlenko, Anna
Interlandi, Matteo
Haynes, Brandon
author_facet Liu, Chunwei
Pavlenko, Anna
Interlandi, Matteo
Haynes, Brandon
contents This paper evaluates the suitability of Apache Arrow, Parquet, and ORC as formats for subsumption in an analytical DBMS. We systematically identify and explore the high-level features that are important to support efficient querying in modern OLAP DBMSs and evaluate the ability of each format to support these features. We find that each format has trade-offs that make it more or less suitable for use as a format in a DBMS and identify opportunities to more holistically co-design a unified in-memory and on-disk data representation. Notably, for certain popular machine learning tasks, none of these formats perform optimally, highlighting significant opportunities for advancing format design. Our hope is that this study can be used as a guide for system developers designing and using these formats, as well as provide the community with directions to pursue for improving these common open formats.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Formats in Analytical DBMSs: Performance Trade-offs and Future Directions
Liu, Chunwei
Pavlenko, Anna
Interlandi, Matteo
Haynes, Brandon
Databases
This paper evaluates the suitability of Apache Arrow, Parquet, and ORC as formats for subsumption in an analytical DBMS. We systematically identify and explore the high-level features that are important to support efficient querying in modern OLAP DBMSs and evaluate the ability of each format to support these features. We find that each format has trade-offs that make it more or less suitable for use as a format in a DBMS and identify opportunities to more holistically co-design a unified in-memory and on-disk data representation. Notably, for certain popular machine learning tasks, none of these formats perform optimally, highlighting significant opportunities for advancing format design. Our hope is that this study can be used as a guide for system developers designing and using these formats, as well as provide the community with directions to pursue for improving these common open formats.
title Data Formats in Analytical DBMSs: Performance Trade-offs and Future Directions
topic Databases
url https://arxiv.org/abs/2411.14331