Floating-Point Data Transformation for Lossless Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jamalidinan, Samirasadat, Cheshmi, Kazem
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912528043343872
author Jamalidinan, Samirasadat
Cheshmi, Kazem
author_facet Jamalidinan, Samirasadat
Cheshmi, Kazem
contents Floating-point data is widely used across various domains. Depending on the required precision, each floating-point value can occupy several bytes. Lossless storage of this information is crucial due to its critical accuracy, as seen in applications such as medical imaging and language model weights. In these cases, data size is often significant, making lossless compression essential. Previous approaches either treat this data as raw byte streams for compression or fail to leverage all patterns within the dataset. However, because multiple bytes represent a single value and due to inherent patterns in floating-point representations, some of these bytes are correlated. To leverage this property, we propose a novel data transformation method called Typed Data Transformation (TDT) that groups related bytes together to improve compression. We implemented and tested our approach on various datasets across both CPU and GPU. TDT achieves a geometric mean compression ratio improvement of 1.16$\times$ over state-of-the-art compression tools such as zstd, while also improving both compression and decompression throughput by 1.18--3.79$\times$.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18062
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Floating-Point Data Transformation for Lossless Compression
Jamalidinan, Samirasadat
Cheshmi, Kazem
Databases
Distributed, Parallel, and Cluster Computing
Floating-point data is widely used across various domains. Depending on the required precision, each floating-point value can occupy several bytes. Lossless storage of this information is crucial due to its critical accuracy, as seen in applications such as medical imaging and language model weights. In these cases, data size is often significant, making lossless compression essential. Previous approaches either treat this data as raw byte streams for compression or fail to leverage all patterns within the dataset. However, because multiple bytes represent a single value and due to inherent patterns in floating-point representations, some of these bytes are correlated. To leverage this property, we propose a novel data transformation method called Typed Data Transformation (TDT) that groups related bytes together to improve compression. We implemented and tested our approach on various datasets across both CPU and GPU. TDT achieves a geometric mean compression ratio improvement of 1.16$\times$ over state-of-the-art compression tools such as zstd, while also improving both compression and decompression throughput by 1.18--3.79$\times$.
title Floating-Point Data Transformation for Lossless Compression
topic Databases
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2506.18062