MeanCache: User-Centric Semantic Caching for LLM Web Services
Fuente:
arXiv
Saved in:
| Main Authors: | Gill, Waris, Elidrisi, Mohamed, Kalapatapu, Pallavi, Ahmed, Ammar, Anwar, Ali, Gulzar, Muhammad Ali |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FedDebug: Systematic Debugging for Federated Learning Applications
by: Gill, Waris, et al.
Published: (2023)
by: Gill, Waris, et al.
Published: (2023)
TraceFL: Interpretability-Driven Debugging in Federated Learning via Neuron Provenance
by: Gill, Waris, et al.
Published: (2023)
by: Gill, Waris, et al.
Published: (2023)
Adaptive K-PackCache: Cost-Centric Data Caching in Cloud
by: Sarkar, Suvarthi, et al.
Published: (2025)
by: Sarkar, Suvarthi, et al.
Published: (2025)
Experimental Analysis of Server-Side Caching for Web Performance
by: Umar, Mohammad, et al.
Published: (2026)
by: Umar, Mohammad, et al.
Published: (2026)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
by: Liu, Kaiwei, et al.
Published: (2025)
by: Liu, Kaiwei, et al.
Published: (2025)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
by: Afroz, Sabiha, et al.
Published: (2025)
by: Afroz, Sabiha, et al.
Published: (2025)
ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
by: Chen, Xinhang, et al.
Published: (2025)
by: Chen, Xinhang, et al.
Published: (2025)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
by: Nian, Sean, et al.
Published: (2026)
by: Nian, Sean, et al.
Published: (2026)
TDC-Cache: A Trustworthy Decentralized Cooperative Caching Framework for Web3.0
by: Chen, Jinyu, et al.
Published: (2025)
by: Chen, Jinyu, et al.
Published: (2025)
FedCache: A Knowledge Cache-driven Federated Learning Architecture for Personalized Edge Intelligence
by: Wu, Zhiyuan, et al.
Published: (2023)
by: Wu, Zhiyuan, et al.
Published: (2023)
InstCache: A Predictive Cache for LLM Serving
by: Zou, Longwei, et al.
Published: (2024)
by: Zou, Longwei, et al.
Published: (2024)
CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
by: Yi, Mengjun, et al.
Published: (2025)
by: Yi, Mengjun, et al.
Published: (2025)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
by: Fang, Shaoke, et al.
Published: (2026)
by: Fang, Shaoke, et al.
Published: (2026)
Caching Aided Multi-Tenant Serverless Computing
by: Qiao, Chu, et al.
Published: (2024)
by: Qiao, Chu, et al.
Published: (2024)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
by: Lee, Sanghyeon, et al.
Published: (2025)
by: Lee, Sanghyeon, et al.
Published: (2025)
LLM-dCache: Improving Tool-Augmented LLMs with GPT-Driven Localized Data Caching
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
by: Zhu, Jianian, et al.
Published: (2025)
by: Zhu, Jianian, et al.
Published: (2025)
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
FLeeC: a Fast Lock-Free Application Cache
by: Costa, André J., et al.
Published: (2024)
by: Costa, André J., et al.
Published: (2024)
KV Cache Compression for Inference Efficiency in LLMs: A Review
by: Liu, Yanyu, et al.
Published: (2025)
by: Liu, Yanyu, et al.
Published: (2025)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
by: Xie, Zhiqiang, et al.
Published: (2025)
by: Xie, Zhiqiang, et al.
Published: (2025)
Comparative Analysis of Distributed Caching Algorithms: Performance Metrics and Implementation Considerations
by: Mayer, Helen, et al.
Published: (2025)
by: Mayer, Helen, et al.
Published: (2025)
CARM Tool: Cache-Aware Roofline Model Automatic Benchmarking and Application Analysis
by: Morgado, José, et al.
Published: (2026)
by: Morgado, José, et al.
Published: (2026)
Bridging Cache-Friendliness and Concurrency: A Locality-Optimized In-Memory B-Skiplist
by: Luo, Yicong, et al.
Published: (2025)
by: Luo, Yicong, et al.
Published: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
by: Zhang, Hanze, et al.
Published: (2025)
by: Zhang, Hanze, et al.
Published: (2025)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
by: Zhao, Zhan, et al.
Published: (2026)
by: Zhao, Zhan, et al.
Published: (2026)
Fine-Grained Vectorized Merge Sorting on RISC-V: From Register to Cache
by: Zhang, Jin, et al.
Published: (2024)
by: Zhang, Jin, et al.
Published: (2024)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
by: Muhammad, Said, et al.
Published: (2025)
by: Muhammad, Said, et al.
Published: (2025)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
by: Yuan, Ying, et al.
Published: (2026)
by: Yuan, Ying, et al.
Published: (2026)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026)
by: Wang, Wenfeng, et al.
Published: (2026)
Maxing Out the SVM: Performance Impact of Memory and Program Cache Sizes in the Agave Validator
by: Vural, Turan, et al.
Published: (2025)
by: Vural, Turan, et al.
Published: (2025)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
by: Sandholm, Thomas, et al.
Published: (2025)
by: Sandholm, Thomas, et al.
Published: (2025)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
A Semantic Quantum Circuit Cache for Scalable and Distributed Quantum-Classical Workflows
by: Tejedor, Mar, et al.
Published: (2026)
by: Tejedor, Mar, et al.
Published: (2026)
Data Caching for Enterprise-Grade Petabyte-Scale OLAP
by: Tang, Chunxu, et al.
Published: (2024)
by: Tang, Chunxu, et al.
Published: (2024)
Idiosyncrasies of Programmable Caching Engines
by: Peixoto, José, et al.
Published: (2026)
by: Peixoto, José, et al.
Published: (2026)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
by: Zheng, Xianzhe, et al.
Published: (2026)
by: Zheng, Xianzhe, et al.
Published: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
by: Yoon, Dongha, et al.
Published: (2025)
by: Yoon, Dongha, et al.
Published: (2025)
CausalMesh: A Formally Verified Causally Consistent Distributed Cache with Support for Client Migration
by: Zhang, Haoran, et al.
Published: (2025)
by: Zhang, Haoran, et al.
Published: (2025)
Similar Items
-
FedDebug: Systematic Debugging for Federated Learning Applications
by: Gill, Waris, et al.
Published: (2023) -
TraceFL: Interpretability-Driven Debugging in Federated Learning via Neuron Provenance
by: Gill, Waris, et al.
Published: (2023) -
Adaptive K-PackCache: Cost-Centric Data Caching in Cloud
by: Sarkar, Suvarthi, et al.
Published: (2025) -
Experimental Analysis of Server-Side Caching for Web Performance
by: Umar, Mohammad, et al.
Published: (2026) -
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
by: Liu, Kaiwei, et al.
Published: (2025)