Lightweight Correlation-Aware Table Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stoian, Mihail, van Renen, Alexander, Kobiolka, Jan, Kuo, Ping-Lin, Grabocka, Josif, Kipf, Andreas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917814208561152
author Stoian, Mihail
van Renen, Alexander
Kobiolka, Jan
Kuo, Ping-Lin
Grabocka, Josif
Kipf, Andreas
author_facet Stoian, Mihail
van Renen, Alexander
Kobiolka, Jan
Kuo, Ping-Lin
Grabocka, Josif
Kipf, Andreas
contents The growing adoption of data lakes for managing relational data necessitates efficient, open storage formats that provide high scan performance and competitive compression ratios. While existing formats achieve fast scans through lightweight encoding techniques, they have reached a plateau in terms of minimizing storage footprint. Recently, correlation-aware compression schemes have been shown to reduce file sizes further. Yet, current approaches either incur significant scan overheads or require manual specification of correlations, limiting their practicability. We present $\texttt{Virtual}$, a framework that integrates seamlessly with existing open formats to automatically leverage data correlations, achieving substantial compression gains while having minimal scan performance overhead. Experiments on data-gov datasets show that $\texttt{Virtual}$ reduces file sizes by up to 40% compared to Apache Parquet.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14066
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lightweight Correlation-Aware Table Compression
Stoian, Mihail
van Renen, Alexander
Kobiolka, Jan
Kuo, Ping-Lin
Grabocka, Josif
Kipf, Andreas
Databases
Information Retrieval
Machine Learning
The growing adoption of data lakes for managing relational data necessitates efficient, open storage formats that provide high scan performance and competitive compression ratios. While existing formats achieve fast scans through lightweight encoding techniques, they have reached a plateau in terms of minimizing storage footprint. Recently, correlation-aware compression schemes have been shown to reduce file sizes further. Yet, current approaches either incur significant scan overheads or require manual specification of correlations, limiting their practicability. We present $\texttt{Virtual}$, a framework that integrates seamlessly with existing open formats to automatically leverage data correlations, achieving substantial compression gains while having minimal scan performance overhead. Experiments on data-gov datasets show that $\texttt{Virtual}$ reduces file sizes by up to 40% compared to Apache Parquet.
title Lightweight Correlation-Aware Table Compression
topic Databases
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2410.14066