Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yubeaton, Patrick, Mahmoud, Tareq, Naga, Shehab, Taheri, Pooria, Xia, Tianhua, George, Arun, Khalil, Yasmein, Zhang, Sai Qian, Joshi, Siddharth, Hegde, Chinmay, Garg, Siddharth
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929695159746560
author Yubeaton, Patrick
Mahmoud, Tareq
Naga, Shehab
Taheri, Pooria
Xia, Tianhua
George, Arun
Khalil, Yasmein
Zhang, Sai Qian
Joshi, Siddharth
Hegde, Chinmay
Garg, Siddharth
author_facet Yubeaton, Patrick
Mahmoud, Tareq
Naga, Shehab
Taheri, Pooria
Xia, Tianhua
George, Arun
Khalil, Yasmein
Zhang, Sai Qian
Joshi, Siddharth
Hegde, Chinmay
Garg, Siddharth
contents As they become more capable, large language models (LLMs) have continued to rapidly increase in size. This has exacerbated the difficulty in running state of the art LLMs on small, edge devices. Standard techniques advocate solving this problem through lossy compression techniques such as quantization or pruning. However, such compression techniques are lossy, and have been shown to change model behavior in unpredictable manners. We propose Huff-LLM, an \emph{end-to-end, lossless} model compression method that lets users store LLM weights in compressed format \emph{everywhere} -- cloud, disk, main memory, and even in on-chip memory/buffers. This allows us to not only load larger models in main memory, but also reduces bandwidth required to load weights on chip, and makes more efficient use of on-chip weight buffers. In addition to the memory savings achieved via compression, we also show latency and energy efficiency improvements when performing inference with the compressed model.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Yubeaton, Patrick
Mahmoud, Tareq
Naga, Shehab
Taheri, Pooria
Xia, Tianhua
George, Arun
Khalil, Yasmein
Zhang, Sai Qian
Joshi, Siddharth
Hegde, Chinmay
Garg, Siddharth
Machine Learning
Hardware Architecture
As they become more capable, large language models (LLMs) have continued to rapidly increase in size. This has exacerbated the difficulty in running state of the art LLMs on small, edge devices. Standard techniques advocate solving this problem through lossy compression techniques such as quantization or pruning. However, such compression techniques are lossy, and have been shown to change model behavior in unpredictable manners. We propose Huff-LLM, an \emph{end-to-end, lossless} model compression method that lets users store LLM weights in compressed format \emph{everywhere} -- cloud, disk, main memory, and even in on-chip memory/buffers. This allows us to not only load larger models in main memory, but also reduces bandwidth required to load weights on chip, and makes more efficient use of on-chip weight buffers. In addition to the memory savings achieved via compression, we also show latency and energy efficiency improvements when performing inference with the compressed model.
title Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
topic Machine Learning
Hardware Architecture
url https://arxiv.org/abs/2502.00922