Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929695159746560 |
|---|---|
| author | Yubeaton, Patrick Mahmoud, Tareq Naga, Shehab Taheri, Pooria Xia, Tianhua George, Arun Khalil, Yasmein Zhang, Sai Qian Joshi, Siddharth Hegde, Chinmay Garg, Siddharth |
| author_facet | Yubeaton, Patrick Mahmoud, Tareq Naga, Shehab Taheri, Pooria Xia, Tianhua George, Arun Khalil, Yasmein Zhang, Sai Qian Joshi, Siddharth Hegde, Chinmay Garg, Siddharth |
| contents | As they become more capable, large language models (LLMs) have continued to rapidly increase in size. This has exacerbated the difficulty in running state of the art LLMs on small, edge devices. Standard techniques advocate solving this problem through lossy compression techniques such as quantization or pruning. However, such compression techniques are lossy, and have been shown to change model behavior in unpredictable manners. We propose Huff-LLM, an \emph{end-to-end, lossless} model compression method that lets users store LLM weights in compressed format \emph{everywhere} -- cloud, disk, main memory, and even in on-chip memory/buffers. This allows us to not only load larger models in main memory, but also reduces bandwidth required to load weights on chip, and makes more efficient use of on-chip weight buffers. In addition to the memory savings achieved via compression, we also show latency and energy efficiency improvements when performing inference with the compressed model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_00922 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Yubeaton, Patrick Mahmoud, Tareq Naga, Shehab Taheri, Pooria Xia, Tianhua George, Arun Khalil, Yasmein Zhang, Sai Qian Joshi, Siddharth Hegde, Chinmay Garg, Siddharth Machine Learning Hardware Architecture As they become more capable, large language models (LLMs) have continued to rapidly increase in size. This has exacerbated the difficulty in running state of the art LLMs on small, edge devices. Standard techniques advocate solving this problem through lossy compression techniques such as quantization or pruning. However, such compression techniques are lossy, and have been shown to change model behavior in unpredictable manners. We propose Huff-LLM, an \emph{end-to-end, lossless} model compression method that lets users store LLM weights in compressed format \emph{everywhere} -- cloud, disk, main memory, and even in on-chip memory/buffers. This allows us to not only load larger models in main memory, but also reduces bandwidth required to load weights on chip, and makes more efficient use of on-chip weight buffers. In addition to the memory savings achieved via compression, we also show latency and energy efficiency improvements when performing inference with the compressed model. |
| title | Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference |
| topic | Machine Learning Hardware Architecture |
| url | https://arxiv.org/abs/2502.00922 |