AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Shaoting, Li, Hanchen, Du, Kuntai, Gu, Zhuohan, Liu, Yuhan, Yao, Jiayi, Ray, Siddhant, Shen, Samuel, Cheng, Yihua, Ananthanarayanan, Ganesh, Jiang, Junchen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023)
by: Liu, Yuhan, et al.
Published: (2023)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
by: Liu, Yuhan, et al.
Published: (2024)
by: Liu, Yuhan, et al.
Published: (2024)
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
by: Yao, Jiayi, et al.
Published: (2024)
by: Yao, Jiayi, et al.
Published: (2024)
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation
by: Ray, Siddhant, et al.
Published: (2024)
by: Ray, Siddhant, et al.
Published: (2024)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
Eloquent: A More Robust Transmission Scheme for LLM Token Streaming
by: Li, Hanchen, et al.
Published: (2024)
by: Li, Hanchen, et al.
Published: (2024)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
2DIO: A Cache-Accurate Storage Microbenchmark
by: Wang, Yirong, et al.
Published: (2026)
by: Wang, Yirong, et al.
Published: (2026)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
by: Wang, Shihao, et al.
Published: (2026)
by: Wang, Shihao, et al.
Published: (2026)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
by: Xiang, Xingyu, et al.
Published: (2025)
by: Xiang, Xingyu, et al.
Published: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Cache is King: Smart Page Eviction with eBPF
by: Zussman, Tal, et al.
Published: (2025)
by: Zussman, Tal, et al.
Published: (2025)
HACache: Leveraging Read Performance with Cache in a Heterogeneous Array
by: Liu, Jialin, et al.
Published: (2026)
by: Liu, Jialin, et al.
Published: (2026)
Getting the MOST out of your Storage Hierarchy with Mirror-Optimized Storage Tiering
by: Tu, Kaiwei, et al.
Published: (2025)
by: Tu, Kaiwei, et al.
Published: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
by: Wu, Fangnuo, et al.
Published: (2025)
by: Wu, Fangnuo, et al.
Published: (2025)
Idiosyncrasies of Programmable Caching Engines
by: Peixoto, José, et al.
Published: (2026)
by: Peixoto, José, et al.
Published: (2026)
DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace File Systems
by: Li, Haoyu, et al.
Published: (2025)
by: Li, Haoyu, et al.
Published: (2025)
Optimizing SSD Caches for Cloud Block Storage Systems Using Machine Learning Approaches
by: Cheng, Chiyu, et al.
Published: (2024)
by: Cheng, Chiyu, et al.
Published: (2024)
HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse
by: An, Yuwei, et al.
Published: (2025)
by: An, Yuwei, et al.
Published: (2025)
DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing
by: Berend, Daniel, et al.
Published: (2025)
by: Berend, Daniel, et al.
Published: (2025)
OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
by: Du, Kuntai, et al.
Published: (2023)
by: Du, Kuntai, et al.
Published: (2023)
DPC: A Distributed Page Cache over CXL
by: Bergman, Shai, et al.
Published: (2026)
by: Bergman, Shai, et al.
Published: (2026)
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
by: Agarwal, Shubham, et al.
Published: (2025)
by: Agarwal, Shubham, et al.
Published: (2025)
Free products and rescalings involving non-separable abelian von Neumann algebras
by: Dykema, Ken, et al.
Published: (2025)
by: Dykema, Ken, et al.
Published: (2025)
LearnedCache: An eBPF-Integrated Perceptron-Based Eviction Policy for the Linux Page Cache
by: Qi, Zejia
Published: (2026)
by: Qi, Zejia
Published: (2026)
I/O Transit Caching for PMem-based Block Device
by: Xu, Qing, et al.
Published: (2024)
by: Xu, Qing, et al.
Published: (2024)
Random Adaptive Cache Placement Policy
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache Resources
by: Kanellopoulos, Konstantinos, et al.
Published: (2023)
by: Kanellopoulos, Konstantinos, et al.
Published: (2023)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
by: Zhang, Dingyan, et al.
Published: (2024)
by: Zhang, Dingyan, et al.
Published: (2024)
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
by: Cai, Tianhao, et al.
Published: (2025)
by: Cai, Tianhao, et al.
Published: (2025)
CRISP: Confidentiality, Rollback, and Integrity Storage Protection for Confidential Cloud-Native Computing
by: Hartono, Ardhi Putra Pratama, et al.
Published: (2024)
by: Hartono, Ardhi Putra Pratama, et al.
Published: (2024)
Similar Items
-
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025) -
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023) -
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
by: Liu, Yuhan, et al.
Published: (2024) -
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
by: Yao, Jiayi, et al.
Published: (2024) -
Towards More Economical Context-Augmented LLM Generation by Reusing Stored KV Cache
by: Li, Hanchen, et al.
Published: (2025)