Lightweight Correlation-Aware Table Compression
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917814208561152 |
|---|---|
| author | Stoian, Mihail van Renen, Alexander Kobiolka, Jan Kuo, Ping-Lin Grabocka, Josif Kipf, Andreas |
| author_facet | Stoian, Mihail van Renen, Alexander Kobiolka, Jan Kuo, Ping-Lin Grabocka, Josif Kipf, Andreas |
| contents | The growing adoption of data lakes for managing relational data necessitates efficient, open storage formats that provide high scan performance and competitive compression ratios. While existing formats achieve fast scans through lightweight encoding techniques, they have reached a plateau in terms of minimizing storage footprint. Recently, correlation-aware compression schemes have been shown to reduce file sizes further. Yet, current approaches either incur significant scan overheads or require manual specification of correlations, limiting their practicability. We present $\texttt{Virtual}$, a framework that integrates seamlessly with existing open formats to automatically leverage data correlations, achieving substantial compression gains while having minimal scan performance overhead. Experiments on data-gov datasets show that $\texttt{Virtual}$ reduces file sizes by up to 40% compared to Apache Parquet. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_14066 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Lightweight Correlation-Aware Table Compression Stoian, Mihail van Renen, Alexander Kobiolka, Jan Kuo, Ping-Lin Grabocka, Josif Kipf, Andreas Databases Information Retrieval Machine Learning The growing adoption of data lakes for managing relational data necessitates efficient, open storage formats that provide high scan performance and competitive compression ratios. While existing formats achieve fast scans through lightweight encoding techniques, they have reached a plateau in terms of minimizing storage footprint. Recently, correlation-aware compression schemes have been shown to reduce file sizes further. Yet, current approaches either incur significant scan overheads or require manual specification of correlations, limiting their practicability. We present $\texttt{Virtual}$, a framework that integrates seamlessly with existing open formats to automatically leverage data correlations, achieving substantial compression gains while having minimal scan performance overhead. Experiments on data-gov datasets show that $\texttt{Virtual}$ reduces file sizes by up to 40% compared to Apache Parquet. |
| title | Lightweight Correlation-Aware Table Compression |
| topic | Databases Information Retrieval Machine Learning |
| url | https://arxiv.org/abs/2410.14066 |