EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Feng, Shaoting, Liu, Yuhan, Li, Hanchen, Chen, Xiaokun, Shen, Samuel, Du, Kuntai, Gu, Zhuohan, Zhang, Rui, Huang, Yuyang, Cheng, Yihua, Yao, Jiayi, Zhang, Qizheng, Ananthanarayanan, Ganesh, Jiang, Junchen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2023)
di: Liu, Yuhan, et al.
Pubblicazione: (2023)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
di: Liu, Yuhan, et al.
Pubblicazione: (2025)
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
di: Yao, Jiayi, et al.
Pubblicazione: (2024)
di: Yao, Jiayi, et al.
Pubblicazione: (2024)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
di: Li, Hanchen, et al.
Pubblicazione: (2025)
di: Li, Hanchen, et al.
Pubblicazione: (2025)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
di: Li, Hanchen, et al.
Pubblicazione: (2025)
di: Li, Hanchen, et al.
Pubblicazione: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
di: Qiu, Shi, et al.
Pubblicazione: (2026)
di: Qiu, Shi, et al.
Pubblicazione: (2026)
OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
di: Du, Kuntai, et al.
Pubblicazione: (2023)
di: Du, Kuntai, et al.
Pubblicazione: (2023)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
di: Li, Ruihao, et al.
Pubblicazione: (2025)
di: Li, Ruihao, et al.
Pubblicazione: (2025)
Cache is King: Smart Page Eviction with eBPF
di: Zussman, Tal, et al.
Pubblicazione: (2025)
di: Zussman, Tal, et al.
Pubblicazione: (2025)
METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation
di: Ray, Siddhant, et al.
Pubblicazione: (2024)
di: Ray, Siddhant, et al.
Pubblicazione: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
di: Gu, Zhuohan, et al.
Pubblicazione: (2024)
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
di: Zhong, Zhiqing, et al.
Pubblicazione: (2026)
di: Zhong, Zhiqing, et al.
Pubblicazione: (2026)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
di: Wang, Shihao, et al.
Pubblicazione: (2026)
di: Wang, Shihao, et al.
Pubblicazione: (2026)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
di: Zou, Jing, et al.
Pubblicazione: (2026)
di: Zou, Jing, et al.
Pubblicazione: (2026)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
di: Chu, Kexin, et al.
Pubblicazione: (2025)
di: Chu, Kexin, et al.
Pubblicazione: (2025)
LearnedCache: An eBPF-Integrated Perceptron-Based Eviction Policy for the Linux Page Cache
di: Qi, Zejia
Pubblicazione: (2026)
di: Qi, Zejia
Pubblicazione: (2026)
Eloquent: A More Robust Transmission Scheme for LLM Token Streaming
di: Li, Hanchen, et al.
Pubblicazione: (2024)
di: Li, Hanchen, et al.
Pubblicazione: (2024)
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
di: Xia, Tian, et al.
Pubblicazione: (2026)
di: Xia, Tian, et al.
Pubblicazione: (2026)
Do Large Language Models Need a Content Delivery Network?
di: Cheng, Yihua, et al.
Pubblicazione: (2024)
di: Cheng, Yihua, et al.
Pubblicazione: (2024)
Free products and rescalings involving non-separable abelian von Neumann algebras
di: Dykema, Ken, et al.
Pubblicazione: (2025)
di: Dykema, Ken, et al.
Pubblicazione: (2025)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
di: Xu, Yuhang, et al.
Pubblicazione: (2025)
di: Xu, Yuhang, et al.
Pubblicazione: (2025)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
di: Wu, Yinpeng, et al.
Pubblicazione: (2026)
di: Wu, Yinpeng, et al.
Pubblicazione: (2026)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
di: Luo, Shutian, et al.
Pubblicazione: (2026)
di: Luo, Shutian, et al.
Pubblicazione: (2026)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
di: Liu, Jinshu, et al.
Pubblicazione: (2024)
di: Liu, Jinshu, et al.
Pubblicazione: (2024)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
di: Chen, Yukang, et al.
Pubblicazione: (2025)
di: Chen, Yukang, et al.
Pubblicazione: (2025)
HACache: Leveraging Read Performance with Cache in a Heterogeneous Array
di: Liu, Jialin, et al.
Pubblicazione: (2026)
di: Liu, Jialin, et al.
Pubblicazione: (2026)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
di: Wu, Fangnuo, et al.
Pubblicazione: (2025)
di: Wu, Fangnuo, et al.
Pubblicazione: (2025)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
di: Zhang, Junkai, et al.
Pubblicazione: (2026)
di: Zhang, Junkai, et al.
Pubblicazione: (2026)
2DIO: A Cache-Accurate Storage Microbenchmark
di: Wang, Yirong, et al.
Pubblicazione: (2026)
di: Wang, Yirong, et al.
Pubblicazione: (2026)
DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace File Systems
di: Li, Haoyu, et al.
Pubblicazione: (2025)
di: Li, Haoyu, et al.
Pubblicazione: (2025)
DPC: A Distributed Page Cache over CXL
di: Bergman, Shai, et al.
Pubblicazione: (2026)
di: Bergman, Shai, et al.
Pubblicazione: (2026)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
A Simple Plug-in for Improving Eviction-Based KV Cache Compression
di: Lin, Yuping, et al.
Pubblicazione: (2026)
di: Lin, Yuping, et al.
Pubblicazione: (2026)
DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing
di: Berend, Daniel, et al.
Pubblicazione: (2025)
di: Berend, Daniel, et al.
Pubblicazione: (2025)
Idiosyncrasies of Programmable Caching Engines
di: Peixoto, José, et al.
Pubblicazione: (2026)
di: Peixoto, José, et al.
Pubblicazione: (2026)
Documenti analoghi
-
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025) -
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2023) -
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024) -
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
di: Liu, Yuhan, et al.
Pubblicazione: (2025) -
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
di: Yao, Jiayi, et al.
Pubblicazione: (2024)