Quad Length Codes for Lossless Compression of e4m3

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Agrawal, Aditya, Magyar, Albert, Eswaraiah, Hiteshwar, Sheridan, Patrick, Janedula, Pradeep, Venkatesan, Ravi Krishnan, Nair, Krishna, Iyer, Ravi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917294273200128
author Agrawal, Aditya
Magyar, Albert
Eswaraiah, Hiteshwar
Sheridan, Patrick
Janedula, Pradeep
Venkatesan, Ravi Krishnan
Nair, Krishna
Iyer, Ravi
author_facet Agrawal, Aditya
Magyar, Albert
Eswaraiah, Hiteshwar
Sheridan, Patrick
Janedula, Pradeep
Venkatesan, Ravi Krishnan
Nair, Krishna
Iyer, Ravi
contents Training and serving Large Language Models (LLMs) relies heavily on parallelization and collective operations, which are frequently bottlenecked by network bandwidth. Lossless compression using e.g., Huffman codes can alleviate the issue, however, Huffman codes suffer from slow, bit-sequential decoding and high hardware complexity due to deep tree traversals. Universal codes e.g., Exponential-Golomb codes are faster to decode but do not exploit the symbol frequency distributions. To address these limitations, this paper introduces Quad Length Codes, a hybrid approach designed to balance compression efficiency with decoding speed. The coding scheme uses 3 prefix bits to divide the 256 symbols into 8 areas. Each area has a different code length and encodes a different number of symbols. The scheme uses a Look Up Table with 256 entries, significantly simplifying the hardware implementation compared to Huffman trees. The coding scheme can be adapted for different distributions. For the e4m3 data type, the scheme achieves a compressibility of 13.9% in comparison to 15.9% achieved by Huffman codes, but it significantly speeds up the decoding and simplifies the hardware complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17849
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quad Length Codes for Lossless Compression of e4m3
Agrawal, Aditya
Magyar, Albert
Eswaraiah, Hiteshwar
Sheridan, Patrick
Janedula, Pradeep
Venkatesan, Ravi Krishnan
Nair, Krishna
Iyer, Ravi
Machine Learning
Information Theory
Training and serving Large Language Models (LLMs) relies heavily on parallelization and collective operations, which are frequently bottlenecked by network bandwidth. Lossless compression using e.g., Huffman codes can alleviate the issue, however, Huffman codes suffer from slow, bit-sequential decoding and high hardware complexity due to deep tree traversals. Universal codes e.g., Exponential-Golomb codes are faster to decode but do not exploit the symbol frequency distributions. To address these limitations, this paper introduces Quad Length Codes, a hybrid approach designed to balance compression efficiency with decoding speed. The coding scheme uses 3 prefix bits to divide the 256 symbols into 8 areas. Each area has a different code length and encodes a different number of symbols. The scheme uses a Look Up Table with 256 entries, significantly simplifying the hardware implementation compared to Huffman trees. The coding scheme can be adapted for different distributions. For the e4m3 data type, the scheme achieves a compressibility of 13.9% in comparison to 15.9% achieved by Huffman codes, but it significantly speeds up the decoding and simplifies the hardware complexity.
title Quad Length Codes for Lossless Compression of e4m3
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2602.17849