KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Bo, Yang, Taolue, Liu, Youyuan, Zhang, Chengming, He, Xubin, Jin, Sian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
Enhancing Lossy Compression Through Cross-Field Information for Scientific Applications
von: Liu, Youyuan, et al.
Veröffentlicht: (2024)
von: Liu, Youyuan, et al.
Veröffentlicht: (2024)
NeurLZ: An Online Neural Learning-Based Method to Enhance Scientific Lossy Compression
von: Jia, Wenqi, et al.
Veröffentlicht: (2024)
von: Jia, Wenqi, et al.
Veröffentlicht: (2024)
GWLZ: A Group-wise Learning-based Lossy Compression Framework for Scientific Data
von: Jia, Wenqi, et al.
Veröffentlicht: (2024)
von: Jia, Wenqi, et al.
Veröffentlicht: (2024)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
von: Zhu, Jianian, et al.
Veröffentlicht: (2025)
von: Zhu, Jianian, et al.
Veröffentlicht: (2025)
STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data
von: Wang, Daoce, et al.
Veröffentlicht: (2025)
von: Wang, Daoce, et al.
Veröffentlicht: (2025)
KV Cache Compression for Inference Efficiency in LLMs: A Review
von: Liu, Yanyu, et al.
Veröffentlicht: (2025)
von: Liu, Yanyu, et al.
Veröffentlicht: (2025)
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
von: Li, Weiqing, et al.
Veröffentlicht: (2025)
von: Li, Weiqing, et al.
Veröffentlicht: (2025)
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography Framework
von: Luo, Xubin, et al.
Veröffentlicht: (2026)
von: Luo, Xubin, et al.
Veröffentlicht: (2026)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
von: Nian, Sean, et al.
Veröffentlicht: (2026)
von: Nian, Sean, et al.
Veröffentlicht: (2026)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
von: Xiang, Xingyu, et al.
Veröffentlicht: (2025)
von: Xiang, Xingyu, et al.
Veröffentlicht: (2025)
A Survey on Error-Bounded Lossy Compression for Scientific Datasets
von: Di, Sheng, et al.
Veröffentlicht: (2024)
von: Di, Sheng, et al.
Veröffentlicht: (2024)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
PiKV: KV Cache Management System for Mixture of Experts
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
An Optimized Error-controlled MPI Collective Framework Integrated with Lossy Compression
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
von: Zhao, Zhan, et al.
Veröffentlicht: (2026)
von: Zhao, Zhan, et al.
Veröffentlicht: (2026)
Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless Orchestration
von: Wu, Shixun, et al.
Veröffentlicht: (2025)
von: Wu, Shixun, et al.
Veröffentlicht: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
von: Yüzügüler, Ahmet Caner, et al.
Veröffentlicht: (2025)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
A Survey on Large Language Model Acceleration based on KV Cache Management
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
IPComp: Interpolation Based Progressive Lossy Compression for Scientific Applications
von: Yang, Zhuoxun, et al.
Veröffentlicht: (2025)
von: Yang, Zhuoxun, et al.
Veröffentlicht: (2025)
To Compress or Not To Compress: Energy Trade-Offs and Benefits of Lossy Compressed I/O
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
Fast Topology-Aware Lossy Data Compression with Full Preservation of Critical Points and Local Order
von: Fallin, Alex, et al.
Veröffentlicht: (2026)
von: Fallin, Alex, et al.
Veröffentlicht: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
von: Bian, Zhuohang, et al.
Veröffentlicht: (2026)
von: Bian, Zhuohang, et al.
Veröffentlicht: (2026)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
von: Zou, Jing, et al.
Veröffentlicht: (2026)
von: Zou, Jing, et al.
Veröffentlicht: (2026)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
Leyline: KV Cache Directives for Agentic Inference
von: Ma, Bole, et al.
Veröffentlicht: (2026)
von: Ma, Bole, et al.
Veröffentlicht: (2026)
cuSZ-$i$: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
von: Liu, Jinyang, et al.
Veröffentlicht: (2023)
von: Liu, Jinyang, et al.
Veröffentlicht: (2023)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
von: li, Fei, et al.
Veröffentlicht: (2026)
von: li, Fei, et al.
Veröffentlicht: (2026)
ZCCL: Significantly Improving Collective Communication With Error-Bounded Lossy Compression
von: Huang, Jiajun, et al.
Veröffentlicht: (2025)
von: Huang, Jiajun, et al.
Veröffentlicht: (2025)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
von: Patel, Ishan, et al.
Veröffentlicht: (2026)
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
von: Liao, Mengqi, et al.
Veröffentlicht: (2026)
von: Liao, Mengqi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
von: Jiang, Bo, et al.
Veröffentlicht: (2025) -
Enhancing Lossy Compression Through Cross-Field Information for Scientific Applications
von: Liu, Youyuan, et al.
Veröffentlicht: (2024) -
NeurLZ: An Online Neural Learning-Based Method to Enhance Scientific Lossy Compression
von: Jia, Wenqi, et al.
Veröffentlicht: (2024) -
GWLZ: A Group-wise Learning-based Lossy Compression Framework for Scientific Data
von: Jia, Wenqi, et al.
Veröffentlicht: (2024) -
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)