CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Fahey, Ryan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913171979108352
author Fahey, Ryan
author_facet Fahey, Ryan
contents Over the past year, prompt caching in Large Language Models (LLMs) has become increasingly more popular across inference APIs. Prompt caching helps save precious compute resources and speeds up response times by reusing parts of the KV cache of a specific prompt for another request. However, many implementations of prompt caching are not secure against timing attacks or even basic metadata disclosure. Gu et al. (ICML 2025) develop a method to audit prompt caching in LLMs. This paper investigates whether OpenRouter's API gateway architecture introduces prompt caching vulnerabilities that bypass provider-level prompt cache isolation guarantees. Most LLM inference providers implement per-account or per-organization prompt caching to prevent data leaks, but does routing through OpenRouter with shared organizational credentials inadvertently create global cache sharing across all OpenRouter users?
format Preprint
id arxiv_https___arxiv_org_abs_2605_30613
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
Fahey, Ryan
Cryptography and Security
Machine Learning
Over the past year, prompt caching in Large Language Models (LLMs) has become increasingly more popular across inference APIs. Prompt caching helps save precious compute resources and speeds up response times by reusing parts of the KV cache of a specific prompt for another request. However, many implementations of prompt caching are not secure against timing attacks or even basic metadata disclosure. Gu et al. (ICML 2025) develop a method to audit prompt caching in LLMs. This paper investigates whether OpenRouter's API gateway architecture introduces prompt caching vulnerabilities that bypass provider-level prompt cache isolation guarantees. Most LLM inference providers implement per-account or per-organization prompt caching to prevent data leaks, but does routing through OpenRouter with shared organizational credentials inadvertently create global cache sharing across all OpenRouter users?
title CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2605.30613