HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zeng, Bowen, Ren, Feiyang, Zhang, Jun, Gu, Xiaoling, Chen, Ke, Shou, Lidan, Li, Huan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
von: Gu, Yifeng, et al.
Veröffentlicht: (2025)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
von: Zhang, Te, et al.
Veröffentlicht: (2025)
von: Zhang, Te, et al.
Veröffentlicht: (2025)
The Pitfalls of KV Cache Compression
von: Chen, Alex, et al.
Veröffentlicht: (2025)
von: Chen, Alex, et al.
Veröffentlicht: (2025)
AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference
von: Huang, Kai, et al.
Veröffentlicht: (2025)
von: Huang, Kai, et al.
Veröffentlicht: (2025)
MiniCache: KV Cache Compression in Depth Dimension for Large Language Models
von: Liu, Akide, et al.
Veröffentlicht: (2024)
von: Liu, Akide, et al.
Veröffentlicht: (2024)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Lossless KV Cache Compression to 2%
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
von: Yang, Zhen, et al.
Veröffentlicht: (2024)
Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy
von: Huang, Yingbing, et al.
Veröffentlicht: (2026)
von: Huang, Yingbing, et al.
Veröffentlicht: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
von: Liu, Hongyao, et al.
Veröffentlicht: (2026)
von: Liu, Hongyao, et al.
Veröffentlicht: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
FairKV: Balancing Per-Head KV Cache for Fast Multi-GPU Inference
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Bingzhe, et al.
Veröffentlicht: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
von: Yang, Yaoxin, et al.
Veröffentlicht: (2025)
von: Yang, Yaoxin, et al.
Veröffentlicht: (2025)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
von: Yang, Bin, et al.
Veröffentlicht: (2025)
von: Yang, Bin, et al.
Veröffentlicht: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
Cross-Self KV Cache Pruning for Efficient Vision-Language Inference
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2024)
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2024)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
Efficient Long-Horizon GUI Agents via Training-Free KV Cache Compression
von: Zhou, Bowen, et al.
Veröffentlicht: (2026)
von: Zhou, Bowen, et al.
Veröffentlicht: (2026)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models
von: Ma, Xindian, et al.
Veröffentlicht: (2026)
von: Ma, Xindian, et al.
Veröffentlicht: (2026)
Palu: Compressing KV-Cache with Low-Rank Projection
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
von: Chang, Chi-Chih, et al.
Veröffentlicht: (2024)
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
von: Godey, Nathan, et al.
Veröffentlicht: (2025)
CommVQ: Commutative Vector Quantization for KV Cache Compression
von: Li, Junyan, et al.
Veröffentlicht: (2025)
von: Li, Junyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
von: Gu, Yifeng, et al.
Veröffentlicht: (2025) -
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025) -
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
von: Cai, Zefan, et al.
Veröffentlicht: (2025) -
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025) -
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)