Learned Data Compression: Challenges and Opportunities for the Future

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Qiyu, Han, Siyuan, Liao, Jianwei, Li, Jin, Peng, Jingshu, Du, Jun, Chen, Lei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917868595052544
author Liu, Qiyu
Han, Siyuan
Liao, Jianwei
Li, Jin
Peng, Jingshu
Du, Jun
Chen, Lei
author_facet Liu, Qiyu
Han, Siyuan
Liao, Jianwei
Li, Jin
Peng, Jingshu
Du, Jun
Chen, Lei
contents Compressing integer keys is a fundamental operation among multiple communities, such as database management (DB), information retrieval (IR), and high-performance computing (HPC). Recent advances in \emph{learned indexes} have inspired the development of \emph{learned compressors}, which leverage simple yet compact machine learning (ML) models to compress large-scale sorted keys. The core idea behind learned compressors is to \emph{losslessly} encode sorted keys by approximating them with \emph{error-bounded} ML models (e.g., piecewise linear functions) and using a \emph{residual array} to guarantee accurate key reconstruction. While the concept of learned compressors remains in its early stages of exploration, our benchmark results demonstrate that an SIMD-optimized learned compressor can significantly outperform state-of-the-art CPU-based compressors. Drawing on our preliminary experiments, this vision paper explores the potential of learned data compression to enhance critical areas in DBMS and related domains. Furthermore, we outline the key technical challenges that existing systems must address when integrating this emerging methodology.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10770
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learned Data Compression: Challenges and Opportunities for the Future
Liu, Qiyu
Han, Siyuan
Liao, Jianwei
Li, Jin
Peng, Jingshu
Du, Jun
Chen, Lei
Databases
Information Retrieval
Compressing integer keys is a fundamental operation among multiple communities, such as database management (DB), information retrieval (IR), and high-performance computing (HPC). Recent advances in \emph{learned indexes} have inspired the development of \emph{learned compressors}, which leverage simple yet compact machine learning (ML) models to compress large-scale sorted keys. The core idea behind learned compressors is to \emph{losslessly} encode sorted keys by approximating them with \emph{error-bounded} ML models (e.g., piecewise linear functions) and using a \emph{residual array} to guarantee accurate key reconstruction. While the concept of learned compressors remains in its early stages of exploration, our benchmark results demonstrate that an SIMD-optimized learned compressor can significantly outperform state-of-the-art CPU-based compressors. Drawing on our preliminary experiments, this vision paper explores the potential of learned data compression to enhance critical areas in DBMS and related domains. Furthermore, we outline the key technical challenges that existing systems must address when integrating this emerging methodology.
title Learned Data Compression: Challenges and Opportunities for the Future
topic Databases
Information Retrieval
url https://arxiv.org/abs/2412.10770