Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Feng, Guo, Cong, Wei, Chiyue, Zhang, Junyao, Zhou, Changchun, Hanson, Edward, Zhang, Jiaqi, Liu, Xiaoxiao, Li, Hai "Helen", Chen, Yiran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915282862211072
author Cheng, Feng
Guo, Cong
Wei, Chiyue
Zhang, Junyao
Zhou, Changchun
Hanson, Edward
Zhang, Jiaqi
Liu, Xiaoxiao
Li, Hai "Helen"
Chen, Yiran
author_facet Cheng, Feng
Guo, Cong
Wei, Chiyue
Zhang, Junyao
Zhou, Changchun
Hanson, Edward
Zhang, Jiaqi
Liu, Xiaoxiao
Li, Hai "Helen"
Chen, Yiran
contents Large language models (LLMs) have demonstrated transformative capabilities across diverse artificial intelligence applications, yet their deployment is hindered by substantial memory and computational demands, especially in resource-constrained environments. Quantization techniques have emerged as a critical solution, reducing data precision to enhance memory and computational efficiency. However, existing methods often suffer from high runtime overheads and potential accuracy degradation. To address these challenges, we propose Ecco, an entropy-based cache compression technique tailored for LLMs. Ecco combines group-wise and non-uniform quantization with pre-defined shared k-means patterns and Huffman coding to exploit the inherent entropy characteristics of LLM cache data. Recognizing the inefficiencies of traditional Huffman coding in terms of parallelism and latency, we introduce a novel parallel Huffman-based decoding process with a multi-stage pipeline design, reducing latency by two orders of magnitude and achieving throughput comparable to GPU L2 caches. Comprehensive evaluations demonstrate that Ecco achieves an up to 2.9$\times$ and 1.9$\times$ speedup over the state-of-the-art AWQ and SmoothQuant framework, 2.4$\times$ over the Olive accelerator, all while increasing memory capacity by nearly 4$\times$ and maintaining state-of-the-art LLM accuracy. These results underscore the effectiveness of our entropy-based cache compression in enhancing LLM performance and efficiency, paving the way for more deployable large-scale AI models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06901
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
Cheng, Feng
Guo, Cong
Wei, Chiyue
Zhang, Junyao
Zhou, Changchun
Hanson, Edward
Zhang, Jiaqi
Liu, Xiaoxiao
Li, Hai "Helen"
Chen, Yiran
Hardware Architecture
Large language models (LLMs) have demonstrated transformative capabilities across diverse artificial intelligence applications, yet their deployment is hindered by substantial memory and computational demands, especially in resource-constrained environments. Quantization techniques have emerged as a critical solution, reducing data precision to enhance memory and computational efficiency. However, existing methods often suffer from high runtime overheads and potential accuracy degradation. To address these challenges, we propose Ecco, an entropy-based cache compression technique tailored for LLMs. Ecco combines group-wise and non-uniform quantization with pre-defined shared k-means patterns and Huffman coding to exploit the inherent entropy characteristics of LLM cache data. Recognizing the inefficiencies of traditional Huffman coding in terms of parallelism and latency, we introduce a novel parallel Huffman-based decoding process with a multi-stage pipeline design, reducing latency by two orders of magnitude and achieving throughput comparable to GPU L2 caches. Comprehensive evaluations demonstrate that Ecco achieves an up to 2.9$\times$ and 1.9$\times$ speedup over the state-of-the-art AWQ and SmoothQuant framework, 2.4$\times$ over the Olive accelerator, all while increasing memory capacity by nearly 4$\times$ and maintaining state-of-the-art LLM accuracy. These results underscore the effectiveness of our entropy-based cache compression in enhancing LLM performance and efficiency, paving the way for more deployable large-scale AI models.
title Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
topic Hardware Architecture
url https://arxiv.org/abs/2505.06901