Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hossain, Elias, Saha, Swayamjit, Roy, Somshubhra, Prasad, Ravi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915763893305344
author Hossain, Elias
Saha, Swayamjit
Roy, Somshubhra
Prasad, Ravi
author_facet Hossain, Elias
Saha, Swayamjit
Roy, Somshubhra
Prasad, Ravi
contents Even when prompts and parameters are secured, transformer language models remain vulnerable because their key-value (KV) cache during inference constitutes an overlooked attack surface. This paper introduces Malicious Token Injection (MTI), a modular framework that systematically perturbs cached key vectors at selected layers and timesteps through controlled magnitude and frequency, using additive Gaussian noise, zeroing, and orthogonal rotations. A theoretical analysis quantifies how these perturbations propagate through attention, linking logit deviations to the Frobenius norm of corruption and softmax Lipschitz dynamics. Empirical results show that MTI significantly alters next-token distributions and downstream task performance across GPT-2 and LLaMA-2/7B, as well as destabilizes retrieval-augmented and agentic reasoning pipelines. These findings identify cache integrity as a critical yet underexplored vulnerability in current LLM deployments, positioning cache corruption as a reproducible and theoretically grounded threat model for future robustness and security research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17098
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
Hossain, Elias
Saha, Swayamjit
Roy, Somshubhra
Prasad, Ravi
Cryptography and Security
Artificial Intelligence
Even when prompts and parameters are secured, transformer language models remain vulnerable because their key-value (KV) cache during inference constitutes an overlooked attack surface. This paper introduces Malicious Token Injection (MTI), a modular framework that systematically perturbs cached key vectors at selected layers and timesteps through controlled magnitude and frequency, using additive Gaussian noise, zeroing, and orthogonal rotations. A theoretical analysis quantifies how these perturbations propagate through attention, linking logit deviations to the Frobenius norm of corruption and softmax Lipschitz dynamics. Empirical results show that MTI significantly alters next-token distributions and downstream task performance across GPT-2 and LLaMA-2/7B, as well as destabilizes retrieval-augmented and agentic reasoning pipelines. These findings identify cache integrity as a critical yet underexplored vulnerability in current LLM deployments, positioning cache corruption as a reproducible and theoretically grounded threat model for future robustness and security research.
title Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.17098