Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Feng, Guo, Cong, Wei, Chiyue, Zhang, Junyao, Zhou, Changchun, Hanson, Edward, Zhang, Jiaqi, Liu, Xiaoxiao, Li, Hai "Helen", Chen, Yiran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
by: Shan, Haoxuan, et al.
Published: (2025)
by: Shan, Haoxuan, et al.
Published: (2025)
Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
by: Fu, Yuzhe, et al.
Published: (2025)
by: Fu, Yuzhe, et al.
Published: (2025)
Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
by: Wei, Chiyue, et al.
Published: (2025)
by: Wei, Chiyue, et al.
Published: (2025)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
by: Guo, Cong, et al.
Published: (2025)
by: Guo, Cong, et al.
Published: (2025)
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
by: Cheng, Feng, et al.
Published: (2025)
by: Cheng, Feng, et al.
Published: (2025)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
by: Hong, Jeongmin, et al.
Published: (2024)
by: Hong, Jeongmin, et al.
Published: (2024)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
by: Li, Jindong, et al.
Published: (2025)
by: Li, Jindong, et al.
Published: (2025)
CAMformer: Associative Memory is All You Need
by: Molom-Ochir, Tergel, et al.
Published: (2025)
by: Molom-Ochir, Tergel, et al.
Published: (2025)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
by: Ko, Younghoon, et al.
Published: (2026)
by: Ko, Younghoon, et al.
Published: (2026)
Improving the Representativeness of Simulation Intervals for the Cache Memory System
by: Bueno, Nicolas, et al.
Published: (2024)
by: Bueno, Nicolas, et al.
Published: (2024)
Pushing the Memory Bandwidth Wall with CXL-enabled Idle I/O Bandwidth Harvesting
by: Kadiyala, Divya Kiran, et al.
Published: (2025)
by: Kadiyala, Divya Kiran, et al.
Published: (2025)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Per-Bank Bandwidth Regulation of Shared Last-Level Cache for Real-Time Systems
by: Sullivan, Connor, et al.
Published: (2024)
by: Sullivan, Connor, et al.
Published: (2024)
An O(m+n)-Space Spatiotemporal Denoising Filter with Cache-Like Memories for Dynamic Vision Sensors
by: Zhao, Qinghang, et al.
Published: (2024)
by: Zhao, Qinghang, et al.
Published: (2024)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
by: Ferry, Corentin, et al.
Published: (2024)
by: Ferry, Corentin, et al.
Published: (2024)
Scaling Routers with In-Package Optics and High-Bandwidth Memories
by: Keslassy, Isaac, et al.
Published: (2026)
by: Keslassy, Isaac, et al.
Published: (2026)
Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators
by: Wang, Ruibao, et al.
Published: (2024)
by: Wang, Ruibao, et al.
Published: (2024)
Exploring DRAM Cache Prefetching for Pooled Memory
by: Tirumalasetty, Chandrahas, et al.
Published: (2024)
by: Tirumalasetty, Chandrahas, et al.
Published: (2024)
Per-Bank Memory Bandwidth Regulation for Predictable and Performant Real-Time System
by: Sullivan, Connor Rudy, et al.
Published: (2026)
by: Sullivan, Connor Rudy, et al.
Published: (2026)
Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
by: He, Siyuan, et al.
Published: (2025)
by: He, Siyuan, et al.
Published: (2025)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
by: Gao, Hanyuan, et al.
Published: (2026)
by: Gao, Hanyuan, et al.
Published: (2026)
In-place Switch: Reprogramming based SLC Cache Design for Hybrid 3D SSDs
by: Yang, Xufeng, et al.
Published: (2024)
by: Yang, Xufeng, et al.
Published: (2024)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
by: Zuepke, Alexander, et al.
Published: (2026)
by: Zuepke, Alexander, et al.
Published: (2026)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
by: Tian, Jiayi, et al.
Published: (2025)
by: Tian, Jiayi, et al.
Published: (2025)
Qplacer: Frequency-Aware Component Placement for Superconducting Quantum Computers
by: Zhang, Junyao, et al.
Published: (2024)
by: Zhang, Junyao, et al.
Published: (2024)
Rhea: a Framework for Fast Design and Validation of RTL Cache-Coherent Memory Subsystems
by: Zoni, Davide, et al.
Published: (2025)
by: Zoni, Davide, et al.
Published: (2025)
CacheSquash: Making caches speculation-aware
by: ElAtali, Hossam, et al.
Published: (2024)
by: ElAtali, Hossam, et al.
Published: (2024)
ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture Model
by: Chen, Hanqiu, et al.
Published: (2024)
by: Chen, Hanqiu, et al.
Published: (2024)
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
by: Wang, Yitu, et al.
Published: (2023)
by: Wang, Yitu, et al.
Published: (2023)
Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication Optimization
by: Tan, Zhanhong, et al.
Published: (2024)
by: Tan, Zhanhong, et al.
Published: (2024)
PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-based Long-Context LLM Inference System
by: Kwon, Hyucksung, et al.
Published: (2024)
by: Kwon, Hyucksung, et al.
Published: (2024)
Be CIM or Be Memory: A Dual-mode-aware DNN Compiler for CIM Accelerators
by: Zhao, Shixin, et al.
Published: (2025)
by: Zhao, Shixin, et al.
Published: (2025)
M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type
by: Hu, Weiming, et al.
Published: (2025)
by: Hu, Weiming, et al.
Published: (2025)
PREFENDER: A Prefetching Defender against Cache Side Channel Attacks as A Pretender
by: Li, Luyi, et al.
Published: (2023)
by: Li, Luyi, et al.
Published: (2023)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
Similar Items
-
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
by: Shan, Haoxuan, et al.
Published: (2025) -
Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
by: Wei, Chiyue, et al.
Published: (2025) -
Prosperity: Accelerating Spiking Neural Networks via Product Sparsity
by: Wei, Chiyue, et al.
Published: (2025) -
FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
by: Fu, Yuzhe, et al.
Published: (2025) -
Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
by: Wei, Chiyue, et al.
Published: (2025)