GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
Fuente:
arXiv
Guardado en:
| Autores principales: | Regmi, Sajal, Pun, Chetan Phakami |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Continuous Semantic Caching for Low-Cost LLM Serving
por: Atalar, Baran, et al.
Publicado: (2026)
por: Atalar, Baran, et al.
Publicado: (2026)
vCache: Verified Semantic Prompt Caching
por: Schroeder, Luis Gaspar, et al.
Publicado: (2025)
por: Schroeder, Luis Gaspar, et al.
Publicado: (2025)
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
por: Liu, Xutong, et al.
Publicado: (2025)
por: Liu, Xutong, et al.
Publicado: (2025)
From Exact Hits to Close Enough: Semantic Caching for LLM Embeddings
por: Biton, Dvir David, et al.
Publicado: (2026)
por: Biton, Dvir David, et al.
Publicado: (2026)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
por: Fang, Shaoke, et al.
Publicado: (2026)
por: Fang, Shaoke, et al.
Publicado: (2026)
Semantics-Aware Caching for Concept Learning
por: Teyou, Louis Mozart Kamdem, et al.
Publicado: (2026)
por: Teyou, Louis Mozart Kamdem, et al.
Publicado: (2026)
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
por: Wang, Chen, et al.
Publicado: (2025)
por: Wang, Chen, et al.
Publicado: (2025)
Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
por: Gill, Waris, et al.
Publicado: (2025)
por: Gill, Waris, et al.
Publicado: (2025)
MeanCache: User-Centric Semantic Caching for LLM Web Services
por: Gill, Waris, et al.
Publicado: (2024)
por: Gill, Waris, et al.
Publicado: (2024)
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems
por: Ghaffari, Shervin, et al.
Publicado: (2025)
por: Ghaffari, Shervin, et al.
Publicado: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
por: Ma, Qingsen, et al.
Publicado: (2025)
por: Ma, Qingsen, et al.
Publicado: (2025)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
por: Liu, Guangda, et al.
Publicado: (2024)
por: Liu, Guangda, et al.
Publicado: (2024)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
por: Fu, Tianyu, et al.
Publicado: (2025)
por: Fu, Tianyu, et al.
Publicado: (2025)
Efficient Prompt Caching via Embedding Similarity
por: Zhu, Hanlin, et al.
Publicado: (2024)
por: Zhu, Hanlin, et al.
Publicado: (2024)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
por: Geng, Yingsheng, et al.
Publicado: (2026)
por: Geng, Yingsheng, et al.
Publicado: (2026)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
por: Song, Dinghong, et al.
Publicado: (2025)
por: Song, Dinghong, et al.
Publicado: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
por: Yao, Jiayi, et al.
Publicado: (2026)
por: Yao, Jiayi, et al.
Publicado: (2026)
MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation
por: Noshad, Ali, et al.
Publicado: (2026)
por: Noshad, Ali, et al.
Publicado: (2026)
IC-Cache: Efficient Large Language Model Serving via In-context Caching
por: Yu, Yifan, et al.
Publicado: (2025)
por: Yu, Yifan, et al.
Publicado: (2025)
Caching Techniques for Reducing the Communication Cost of Federated Learning in IoT Environments
por: Alhonainy, Ahmad, et al.
Publicado: (2025)
por: Alhonainy, Ahmad, et al.
Publicado: (2025)
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
por: Liu, Zhiyuan, et al.
Publicado: (2025)
por: Liu, Zhiyuan, et al.
Publicado: (2025)
10Cache: Heterogeneous Resource-Aware Tensor Caching and Migration for LLM Training
por: Afroz, Sabiha, et al.
Publicado: (2025)
por: Afroz, Sabiha, et al.
Publicado: (2025)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
por: Kim, Kihyun, et al.
Publicado: (2025)
por: Kim, Kihyun, et al.
Publicado: (2025)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
Learning-to-Cache: Accelerating Diffusion Transformer via Layer Caching
por: Ma, Xinyin, et al.
Publicado: (2024)
por: Ma, Xinyin, et al.
Publicado: (2024)
User Intent Recognition and Semantic Cache Optimization-Based Query Processing Framework using CFLIS and MGR-LAU
por: Mahendru, Sakshi
Publicado: (2024)
por: Mahendru, Sakshi
Publicado: (2024)
FlexCache: Flexible Approximate Cache System for Video Diffusion
por: Sun, Desen, et al.
Publicado: (2024)
por: Sun, Desen, et al.
Publicado: (2024)
CacheFormer: High Attention-Based Segment Caching
por: Singh, Sushant, et al.
Publicado: (2025)
por: Singh, Sushant, et al.
Publicado: (2025)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
por: Dong, Yanhao, et al.
Publicado: (2025)
por: Dong, Yanhao, et al.
Publicado: (2025)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
por: Couturier, Camille, et al.
Publicado: (2025)
por: Couturier, Camille, et al.
Publicado: (2025)
AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization
por: Tan, Yifan, et al.
Publicado: (2024)
por: Tan, Yifan, et al.
Publicado: (2024)
MEPIC: Memory Efficient Position Independent Caching for LLM Serving
por: Wang, Qian, et al.
Publicado: (2025)
por: Wang, Qian, et al.
Publicado: (2025)
Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
por: Gopal, Nikhil, et al.
Publicado: (2026)
por: Gopal, Nikhil, et al.
Publicado: (2026)
SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
por: Haghighi, Yasaman, et al.
Publicado: (2026)
por: Haghighi, Yasaman, et al.
Publicado: (2026)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
por: Yang, Bin, et al.
Publicado: (2025)
por: Yang, Bin, et al.
Publicado: (2025)
CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
por: Fahey, Ryan
Publicado: (2026)
por: Fahey, Ryan
Publicado: (2026)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
por: Brandon, William, et al.
Publicado: (2024)
por: Brandon, William, et al.
Publicado: (2024)
Sparse Prefix Caching for Hybrid and Recurrent LLM Serving
por: Shirokikh, Mikhail, et al.
Publicado: (2026)
por: Shirokikh, Mikhail, et al.
Publicado: (2026)
Ejemplares similares
-
Continuous Semantic Caching for Low-Cost LLM Serving
por: Atalar, Baran, et al.
Publicado: (2026) -
vCache: Verified Semantic Prompt Caching
por: Schroeder, Luis Gaspar, et al.
Publicado: (2025) -
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
por: Liu, Xutong, et al.
Publicado: (2025) -
From Exact Hits to Close Enough: Semantic Caching for LLM Embeddings
por: Biton, Dvir David, et al.
Publicado: (2026) -
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
por: Fang, Shaoke, et al.
Publicado: (2026)