vCache: Verified Semantic Prompt Caching
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Schroeder, Luis Gaspar, Desai, Aditya, Cuadron, Alejandro, Chu, Kyle, Liu, Shu, Zhao, Mark, Krusche, Stephan, Kemper, Alfons, Zaharia, Matei, Gonzalez, Joseph E. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
vAttention: Verified Sparse Attention
par: Desai, Aditya, et autres
Publié: (2025)
par: Desai, Aditya, et autres
Publié: (2025)
HashAttention: Semantic Sparsity for Faster Inference
par: Desai, Aditya, et autres
Publié: (2024)
par: Desai, Aditya, et autres
Publié: (2024)
Asynchronous Verified Semantic Caching for Tiered LLM Architectures
par: Singh, Asmit Kumar, et autres
Publié: (2026)
par: Singh, Asmit Kumar, et autres
Publié: (2026)
CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
par: Fahey, Ryan
Publié: (2026)
par: Fahey, Ryan
Publié: (2026)
Optimizing LLM Queries in Relational Data Analytics Workloads
par: Liu, Shu, et autres
Publié: (2024)
par: Liu, Shu, et autres
Publié: (2024)
GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
par: Regmi, Sajal, et autres
Publié: (2024)
par: Regmi, Sajal, et autres
Publié: (2024)
MeanCache: User-Centric Semantic Caching for LLM Web Services
par: Gill, Waris, et autres
Publié: (2024)
par: Gill, Waris, et autres
Publié: (2024)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
par: Fu, Tianyu, et autres
Publié: (2025)
par: Fu, Tianyu, et autres
Publié: (2025)
Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
par: Lumer, Elias, et autres
Publié: (2026)
par: Lumer, Elias, et autres
Publié: (2026)
Coded Caching with Shared Caches and Private Caches
par: Peter, Elizabath, et autres
Publié: (2022)
par: Peter, Elizabath, et autres
Publié: (2022)
The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
par: Cuadron, Alejandro, et autres
Publié: (2025)
par: Cuadron, Alejandro, et autres
Publié: (2025)
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
par: Zheng, Zihao, et autres
Publié: (2026)
par: Zheng, Zihao, et autres
Publié: (2026)
SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging
par: Wu, Shunlong, et autres
Publié: (2026)
par: Wu, Shunlong, et autres
Publié: (2026)
Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches
par: Fang, Shaoke, et autres
Publié: (2026)
par: Fang, Shaoke, et autres
Publié: (2026)
Semantic Caching for Improving Web Affordability
par: Akbar, Hafsa, et autres
Publié: (2025)
par: Akbar, Hafsa, et autres
Publié: (2025)
Semantics-Aware Caching for Concept Learning
par: Teyou, Louis Mozart Kamdem, et autres
Publié: (2026)
par: Teyou, Louis Mozart Kamdem, et autres
Publié: (2026)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
par: Tian, Yuyang, et autres
Publié: (2025)
par: Tian, Yuyang, et autres
Publié: (2025)
The Pitfalls of KV Cache Compression
par: Chen, Alex, et autres
Publié: (2025)
par: Chen, Alex, et autres
Publié: (2025)
Generative Caching for Structurally Similar Prompts and Responses
par: Chakraborty, Sarthak, et autres
Publié: (2025)
par: Chakraborty, Sarthak, et autres
Publié: (2025)
Auditing Prompt Caching in Language Model APIs
par: Gu, Chenchen, et autres
Publié: (2025)
par: Gu, Chenchen, et autres
Publié: (2025)
Efficient Prompt Caching via Embedding Similarity
par: Zhu, Hanlin, et autres
Publié: (2024)
par: Zhu, Hanlin, et autres
Publié: (2024)
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
par: Yan, Jianxin, et autres
Publié: (2025)
par: Yan, Jianxin, et autres
Publié: (2025)
OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models
par: Chu, Huanpeng, et autres
Publié: (2025)
par: Chu, Huanpeng, et autres
Publié: (2025)
LLMs for Test Input Generation for Semantic Caches
par: Rasool, Zafaryab, et autres
Publié: (2024)
par: Rasool, Zafaryab, et autres
Publié: (2024)
Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
par: Agarwal, Shubham, et autres
Publié: (2026)
par: Agarwal, Shubham, et autres
Publié: (2026)
Prompt Injection Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching
par: Gosmar, Diego, et autres
Publié: (2026)
par: Gosmar, Diego, et autres
Publié: (2026)
MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation
par: Noshad, Ali, et autres
Publié: (2026)
par: Noshad, Ali, et autres
Publié: (2026)
SIEVE: Sample-Efficient Parametric Learning from Natural Language
par: Asawa, Parth, et autres
Publié: (2026)
par: Asawa, Parth, et autres
Publié: (2026)
PromptTea: Let Prompts Tell TeaCache the Optimal Threshold
par: Huang, Zishen, et autres
Publié: (2025)
par: Huang, Zishen, et autres
Publié: (2025)
Probing the Prompt KV Cache: Where It Becomes Dispensable
par: Kumar, Vinayshekhar Bannihatti, et autres
Publié: (2026)
par: Kumar, Vinayshekhar Bannihatti, et autres
Publié: (2026)
Finch: Prompt-guided Key-Value Cache Compression
par: Corallo, Giulio, et autres
Publié: (2024)
par: Corallo, Giulio, et autres
Publié: (2024)
Alto: Orchestrating Distributed Compound AI Systems with Nested Ancestry
par: Raghavan, Deepti, et autres
Publié: (2024)
par: Raghavan, Deepti, et autres
Publié: (2024)
The Duck's Brain: Training and Inference of Neural Networks in Modern Database Engines
par: Schüle, Maximilian E., et autres
Publié: (2023)
par: Schüle, Maximilian E., et autres
Publié: (2023)
BackCache: Mitigating Contention-Based Cache Timing Attacks by Hiding Cache Line Evictions
par: Wang, Quancheng, et autres
Publié: (2023)
par: Wang, Quancheng, et autres
Publié: (2023)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
par: Basu, Abhinaba
Publié: (2026)
par: Basu, Abhinaba
Publié: (2026)
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
par: Wang, Chen, et autres
Publié: (2025)
par: Wang, Chen, et autres
Publié: (2025)
Continuous Semantic Caching for Low-Cost LLM Serving
par: Atalar, Baran, et autres
Publié: (2026)
par: Atalar, Baran, et autres
Publié: (2026)
A Joint Learning Approach to Hardware Caching and Prefetching
par: Yuan, Samuel, et autres
Publié: (2025)
par: Yuan, Samuel, et autres
Publié: (2025)
InstCache: A Predictive Cache for LLM Serving
par: Zou, Longwei, et autres
Publié: (2024)
par: Zou, Longwei, et autres
Publié: (2024)
Demand Private Coded Caching: Small Cache Size
par: Lu, Qinyi, et autres
Publié: (2025)
par: Lu, Qinyi, et autres
Publié: (2025)
Documents similaires
-
vAttention: Verified Sparse Attention
par: Desai, Aditya, et autres
Publié: (2025) -
HashAttention: Semantic Sparsity for Faster Inference
par: Desai, Aditya, et autres
Publié: (2024) -
Asynchronous Verified Semantic Caching for Tiered LLM Architectures
par: Singh, Asmit Kumar, et autres
Publié: (2026) -
CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
par: Fahey, Ryan
Publié: (2026) -
Optimizing LLM Queries in Relational Data Analytics Workloads
par: Liu, Shu, et autres
Publié: (2024)