CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nahian, Mohaiminul Al, Almalky, Abeer Matar A., Aragonda, Gamana, Zhou, Ranyang, Ahmed, Sabbir, Ponomarev, Dmitry, Yang, Li, Angizi, Shaahin, Rakin, Adnan Siraj
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914508694355968
author Nahian, Mohaiminul Al
Almalky, Abeer Matar A.
Aragonda, Gamana
Zhou, Ranyang
Ahmed, Sabbir
Ponomarev, Dmitry
Yang, Li
Angizi, Shaahin
Rakin, Adnan Siraj
author_facet Nahian, Mohaiminul Al
Almalky, Abeer Matar A.
Aragonda, Gamana
Zhou, Ranyang
Ahmed, Sabbir
Ponomarev, Dmitry
Yang, Li
Angizi, Shaahin
Rakin, Adnan Siraj
contents The rapid advancement of large language models (LLMs) has sparked growing interest in understanding their security vulnerabilities, particularly Trojan attacks that enable stealthy manipulation of model behavior. Traditional Trojan methods typically alter inputs and/or model weights, relying on white-box assumptions that require access to data or model internal parameters. In this work, we present CacheTrap, the first gray-box Trojan attack targeting the Key-Value (KV) cache of LLMs. This method induces a single-bit flip in the KV cache, serving as a transient trigger. When activated, this trigger causes the model to exhibit targeted actions without changing inputs or model weights. CacheTrap introduces an efficient search algorithm to locate vulnerable positions in the KV cache, independent of model weights or datasets. Extensive experiments on five open-source LLMs show a remarkable 100% attack success rate (with the trigger) while preserving benign accuracy (without the trigger) by flipping just one bit in the KV cache.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs
Nahian, Mohaiminul Al
Almalky, Abeer Matar A.
Aragonda, Gamana
Zhou, Ranyang
Ahmed, Sabbir
Ponomarev, Dmitry
Yang, Li
Angizi, Shaahin
Rakin, Adnan Siraj
Cryptography and Security
The rapid advancement of large language models (LLMs) has sparked growing interest in understanding their security vulnerabilities, particularly Trojan attacks that enable stealthy manipulation of model behavior. Traditional Trojan methods typically alter inputs and/or model weights, relying on white-box assumptions that require access to data or model internal parameters. In this work, we present CacheTrap, the first gray-box Trojan attack targeting the Key-Value (KV) cache of LLMs. This method induces a single-bit flip in the KV cache, serving as a transient trigger. When activated, this trigger causes the model to exhibit targeted actions without changing inputs or model weights. CacheTrap introduces an efficient search algorithm to locate vulnerable positions in the KV cache, independent of model weights or datasets. Extensive experiments on five open-source LLMs show a remarkable 100% attack success rate (with the trigger) while preserving benign accuracy (without the trigger) by flipping just one bit in the KV cache.
title CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs
topic Cryptography and Security
url https://arxiv.org/abs/2511.22681