LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yuhan, Cheng, Yihua, Yao, Jiayi, An, Yuwei, Chen, Xiaokun, Feng, Shaoting, Huang, Yuyang, Shen, Samuel, Zhang, Rui, Du, Kuntai, Jiang, Junchen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
di: Li, Hanchen, et al.
Pubblicazione: (2025)
di: Li, Hanchen, et al.
Pubblicazione: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
Do Large Language Models Need a Content Delivery Network?
di: Cheng, Yihua, et al.
Pubblicazione: (2024)
di: Cheng, Yihua, et al.
Pubblicazione: (2024)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2023)
di: Liu, Yuhan, et al.
Pubblicazione: (2023)
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
di: Yao, Jiayi, et al.
Pubblicazione: (2024)
di: Yao, Jiayi, et al.
Pubblicazione: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse
di: An, Yuwei, et al.
Pubblicazione: (2025)
di: An, Yuwei, et al.
Pubblicazione: (2025)
Eloquent: A More Robust Transmission Scheme for LLM Token Streaming
di: Li, Hanchen, et al.
Pubblicazione: (2024)
di: Li, Hanchen, et al.
Pubblicazione: (2024)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
Earth+: on-board satellite imagery compression leveraging historical earth observations
di: Du, Kuntai, et al.
Pubblicazione: (2024)
di: Du, Kuntai, et al.
Pubblicazione: (2024)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
di: Du, Kuntai, et al.
Pubblicazione: (2025)
di: Du, Kuntai, et al.
Pubblicazione: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2026)
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2024)
di: Feng, Yuan, et al.
Pubblicazione: (2024)
METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation
di: Ray, Siddhant, et al.
Pubblicazione: (2024)
di: Ray, Siddhant, et al.
Pubblicazione: (2024)
CachePrune: Privacy-Aware and Fine-Grained KV Cache Sharing for Efficient LLM Inference
di: Wu, Guanlong, et al.
Pubblicazione: (2026)
di: Wu, Guanlong, et al.
Pubblicazione: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
di: Hu, Jie, et al.
Pubblicazione: (2025)
di: Hu, Jie, et al.
Pubblicazione: (2025)
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
di: Zhang, Qiuyang, et al.
Pubblicazione: (2026)
di: Zhang, Qiuyang, et al.
Pubblicazione: (2026)
Taming the Fragility of KV Cache Eviction in LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
di: Xu, Yichun, et al.
Pubblicazione: (2026)
di: Xu, Yichun, et al.
Pubblicazione: (2026)
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
di: Wu, Haoyi, et al.
Pubblicazione: (2024)
di: Wu, Haoyi, et al.
Pubblicazione: (2024)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
di: Zhu, Yuxuan, et al.
Pubblicazione: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
di: Tian, Yuxuan, et al.
Pubblicazione: (2025)
Competitive Non-Clairvoyant KV-Cache Scheduling for LLM Inference
di: Feng, Yiding, et al.
Pubblicazione: (2026)
di: Feng, Yiding, et al.
Pubblicazione: (2026)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
di: Yu, Bohan, et al.
Pubblicazione: (2025)
di: Yu, Bohan, et al.
Pubblicazione: (2025)
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
di: Sharma, Akshat, et al.
Pubblicazione: (2024)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
di: Li, Xing, et al.
Pubblicazione: (2025)
di: Li, Xing, et al.
Pubblicazione: (2025)
Mitigating KV Cache Competition to Enhance User Experience in LLM Inference
di: Shen, Haiying, et al.
Pubblicazione: (2025)
di: Shen, Haiying, et al.
Pubblicazione: (2025)
OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
di: Du, Kuntai, et al.
Pubblicazione: (2023)
di: Du, Kuntai, et al.
Pubblicazione: (2023)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
di: Zhao, Yi, et al.
Pubblicazione: (2025)
di: Zhao, Yi, et al.
Pubblicazione: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
di: Yang, Yifei, et al.
Pubblicazione: (2024)
di: Yang, Yifei, et al.
Pubblicazione: (2024)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
di: Ma, Da, et al.
Pubblicazione: (2024)
di: Ma, Da, et al.
Pubblicazione: (2024)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
di: Chu, Kexin, et al.
Pubblicazione: (2025)
di: Chu, Kexin, et al.
Pubblicazione: (2025)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
di: Zuo, Youhui, et al.
Pubblicazione: (2025)
di: Zuo, Youhui, et al.
Pubblicazione: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026) -
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025) -
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024) -
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
di: Li, Hanchen, et al.
Pubblicazione: (2025) -
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)