From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhixiang, Liu, Zesen, Xie, Yuchong, Huang, Quanfeng, She, Dongdong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914295589109760
author Zhang, Zhixiang
Liu, Zesen
Xie, Yuchong
Huang, Quanfeng
She, Dongdong
author_facet Zhang, Zhixiang
Liu, Zesen
Xie, Yuchong
Huang, Quanfeng
She, Dongdong
contents Semantic caching has emerged as a pivotal technique for scaling LLM applications, widely adopted by major providers including AWS and Microsoft. By utilizing semantic embedding vectors as cache keys, this mechanism effectively minimizes latency and redundant computation for semantically similar queries. In this work, we conceptualize semantic cache keys as a form of fuzzy hashes. We demonstrate that the locality required to maximize cache hit rates fundamentally conflicts with the cryptographic avalanche effect necessary for collision resistance. Our conceptual analysis formalizes this inherent trade-off between performance (locality) and security (collision resilience), revealing that semantic caching is naturally vulnerable to key collision attacks. While prior research has focused on side-channel and privacy risks, we present the first systematic study of integrity risks arising from cache collisions. We introduce CacheAttack, an automated framework for launching black-box collision attacks. We evaluate CacheAttack in security-critical tasks and agentic workflows. It achieves a hit rate of 86\% in LLM response hijacking and can induce malicious behaviors in LLM agent, while preserving strong transferability across different embedding models. A case study on a financial agent further illustrates the real-world impact of these vulnerabilities. Finally, we discuss mitigation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2601_23088
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching
Zhang, Zhixiang
Liu, Zesen
Xie, Yuchong
Huang, Quanfeng
She, Dongdong
Cryptography and Security
Artificial Intelligence
Semantic caching has emerged as a pivotal technique for scaling LLM applications, widely adopted by major providers including AWS and Microsoft. By utilizing semantic embedding vectors as cache keys, this mechanism effectively minimizes latency and redundant computation for semantically similar queries. In this work, we conceptualize semantic cache keys as a form of fuzzy hashes. We demonstrate that the locality required to maximize cache hit rates fundamentally conflicts with the cryptographic avalanche effect necessary for collision resistance. Our conceptual analysis formalizes this inherent trade-off between performance (locality) and security (collision resilience), revealing that semantic caching is naturally vulnerable to key collision attacks. While prior research has focused on side-channel and privacy risks, we present the first systematic study of integrity risks arising from cache collisions. We introduce CacheAttack, an automated framework for launching black-box collision attacks. We evaluate CacheAttack in security-critical tasks and agentic workflows. It achieves a hit rate of 86\% in LLM response hijacking and can induce malicious behaviors in LLM agent, while preserving strong transferability across different embedding models. A case study on a financial agent further illustrates the real-world impact of these vulnerabilities. Finally, we discuss mitigation strategies.
title From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2601.23088