ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Guangda, Li, Chengwei, Zhao, Jieru, Zhang, Chenqi, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
von: Barad, Haim, et al.
Veröffentlicht: (2023)
von: Barad, Haim, et al.
Veröffentlicht: (2023)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
von: Liu, Minghui, et al.
Veröffentlicht: (2025)
von: Liu, Minghui, et al.
Veröffentlicht: (2025)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
von: Liu, Hongyao, et al.
Veröffentlicht: (2026)
von: Liu, Hongyao, et al.
Veröffentlicht: (2026)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
von: Sai, Vishnu, et al.
Veröffentlicht: (2026)
von: Sai, Vishnu, et al.
Veröffentlicht: (2026)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
von: Taneja, Maanas, et al.
Veröffentlicht: (2026)
von: Taneja, Maanas, et al.
Veröffentlicht: (2026)
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
von: Liu, Andy Zeyi, et al.
Veröffentlicht: (2026)
von: Liu, Andy Zeyi, et al.
Veröffentlicht: (2026)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
von: Liu, Minghui, et al.
Veröffentlicht: (2024)
von: Liu, Minghui, et al.
Veröffentlicht: (2024)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
von: Ning, Zhenyu, et al.
Veröffentlicht: (2024)
von: Ning, Zhenyu, et al.
Veröffentlicht: (2024)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
von: Zhang, Te, et al.
Veröffentlicht: (2025)
von: Zhang, Te, et al.
Veröffentlicht: (2025)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
von: Chen, Chuangtao, et al.
Veröffentlicht: (2026)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
von: Yi, Qingao, et al.
Veröffentlicht: (2025)
von: Yi, Qingao, et al.
Veröffentlicht: (2025)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
von: Liu, Sihao, et al.
Veröffentlicht: (2026)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
von: Li, Yubo, et al.
Veröffentlicht: (2026)
von: Li, Yubo, et al.
Veröffentlicht: (2026)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
von: Saxena, Utkarsh, et al.
Veröffentlicht: (2024)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
ManifoldKV: Training-Free KV Cache Compression via Euclidean Outlier Detection
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
von: Datta, Debajyoti, et al.
Veröffentlicht: (2026)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025) -
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025) -
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024) -
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
von: Barad, Haim, et al.
Veröffentlicht: (2023) -
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
von: Liu, Minghui, et al.
Veröffentlicht: (2025)