KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Jiahao, Han, Jinbo, Wei, Xingda, Shen, Sijie, Zhang, Dingyan, Fang, Chenguang, Chen, Rong, Yu, Wenyuan, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
par: Zhang, Dingyan, et autres
Publié: (2026)
par: Zhang, Dingyan, et autres
Publié: (2026)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
par: Zhang, Dingyan, et autres
Publié: (2024)
par: Zhang, Dingyan, et autres
Publié: (2024)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
par: Qin, Ruoyu, et autres
Publié: (2026)
par: Qin, Ruoyu, et autres
Publié: (2026)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
par: Lin, Bin, et autres
Publié: (2024)
par: Lin, Bin, et autres
Publié: (2024)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
par: Qin, Ruoyu, et autres
Publié: (2024)
par: Qin, Ruoyu, et autres
Publié: (2024)
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
par: Cheng, Rongxin, et autres
Publié: (2024)
par: Cheng, Rongxin, et autres
Publié: (2024)
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
par: Yang, Xinjun, et autres
Publié: (2025)
par: Yang, Xinjun, et autres
Publié: (2025)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
par: Nian, Sean, et autres
Publié: (2026)
par: Nian, Sean, et autres
Publié: (2026)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Adaptive K-PackCache: Cost-Centric Data Caching in Cloud
par: Sarkar, Suvarthi, et autres
Publié: (2025)
par: Sarkar, Suvarthi, et autres
Publié: (2025)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
par: Song, Xiaoniu, et autres
Publié: (2024)
par: Song, Xiaoniu, et autres
Publié: (2024)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
par: Cheng, Rongxin, et autres
Publié: (2024)
par: Cheng, Rongxin, et autres
Publié: (2024)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
par: Tofigh, Mani, et autres
Publié: (2025)
par: Tofigh, Mani, et autres
Publié: (2025)
CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
par: Yi, Mengjun, et autres
Publié: (2025)
par: Yi, Mengjun, et autres
Publié: (2025)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
par: Liu, Kaiwei, et autres
Publié: (2025)
par: Liu, Kaiwei, et autres
Publié: (2025)
TDC-Cache: A Trustworthy Decentralized Cooperative Caching Framework for Web3.0
par: Chen, Jinyu, et autres
Publié: (2025)
par: Chen, Jinyu, et autres
Publié: (2025)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
par: Zhu, Jianian, et autres
Publié: (2025)
par: Zhu, Jianian, et autres
Publié: (2025)
FedCache: A Knowledge Cache-driven Federated Learning Architecture for Personalized Edge Intelligence
par: Wu, Zhiyuan, et autres
Publié: (2023)
par: Wu, Zhiyuan, et autres
Publié: (2023)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
par: Wei, Xingda, et autres
Publié: (2024)
par: Wei, Xingda, et autres
Publié: (2024)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
par: Qianli, Liu, et autres
Publié: (2025)
par: Qianli, Liu, et autres
Publié: (2025)
Bridging Cache-Friendliness and Concurrency: A Locality-Optimized In-Memory B-Skiplist
par: Luo, Yicong, et autres
Publié: (2025)
par: Luo, Yicong, et autres
Publié: (2025)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
par: Fang, Shaoke, et autres
Publié: (2026)
par: Fang, Shaoke, et autres
Publié: (2026)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
par: Afroz, Sabiha, et autres
Publié: (2025)
par: Afroz, Sabiha, et autres
Publié: (2025)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
par: Wang, Tianze, et autres
Publié: (2025)
par: Wang, Tianze, et autres
Publié: (2025)
Caching Aided Multi-Tenant Serverless Computing
par: Qiao, Chu, et autres
Publié: (2024)
par: Qiao, Chu, et autres
Publié: (2024)
A Survey on Large Language Model Acceleration based on KV Cache Management
par: Li, Haoyang, et autres
Publié: (2024)
par: Li, Haoyang, et autres
Publié: (2024)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
par: Sandholm, Thomas, et autres
Publié: (2025)
par: Sandholm, Thomas, et autres
Publié: (2025)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
par: Tian, Yuyang, et autres
Publié: (2025)
par: Tian, Yuyang, et autres
Publié: (2025)
InstCache: A Predictive Cache for LLM Serving
par: Zou, Longwei, et autres
Publié: (2024)
par: Zou, Longwei, et autres
Publié: (2024)
Towards Lock Modularization for Heterogeneous Environments
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Data Caching for Enterprise-Grade Petabyte-Scale OLAP
par: Tang, Chunxu, et autres
Publié: (2024)
par: Tang, Chunxu, et autres
Publié: (2024)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
par: Yuan, Ying, et autres
Publié: (2026)
par: Yuan, Ying, et autres
Publié: (2026)
Experimental Analysis of Server-Side Caching for Web Performance
par: Umar, Mohammad, et autres
Publié: (2026)
par: Umar, Mohammad, et autres
Publié: (2026)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
par: Lee, Sanghyeon, et autres
Publié: (2025)
par: Lee, Sanghyeon, et autres
Publié: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
par: Chen, Xinhang, et autres
Publié: (2025)
par: Chen, Xinhang, et autres
Publié: (2025)
FLeeC: a Fast Lock-Free Application Cache
par: Costa, André J., et autres
Publié: (2024)
par: Costa, André J., et autres
Publié: (2024)
KV Cache Compression for Inference Efficiency in LLMs: A Review
par: Liu, Yanyu, et autres
Publié: (2025)
par: Liu, Yanyu, et autres
Publié: (2025)
Documents similaires
-
LMetric: Simple is Better - Multiplication May Be All You Need for LLM Request Scheduling
par: Zhang, Dingyan, et autres
Publié: (2026) -
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
par: Zhang, Dingyan, et autres
Publié: (2024) -
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
par: Zhang, Hanze, et autres
Publié: (2025) -
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
par: Qin, Ruoyu, et autres
Publié: (2026) -
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
par: Lin, Bin, et autres
Publié: (2024)