Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913139900022784 |
|---|---|
| author | Pulipaka, Sidharth Hlebik, Stanislau Raghav, Leonidas Abdelnabi, Sahar Raina, Vyas Sheth, Ivaxi Fritz, Mario |
| author_facet | Pulipaka, Sidharth Hlebik, Stanislau Raghav, Leonidas Abdelnabi, Sahar Raina, Vyas Sheth, Ivaxi Fritz, Mario |
| contents | Large language models are increasingly augmented with persistent memory, allowing assistants to store user-specific information across sessions for personalization and continuity. This statefulness introduces a new security risk: adversarial content can corrupt what an assistant remembers and thereby influence future interactions. We propose and study sleeper memory poisoning, a delayed attack in which an adversary manipulates external context, such as a document, webpage, or repository, to cause the assistant to store a fabricated memory about the user. Unlike conventional prompt injection, the attack can remain dormant and re-emerge across multiple later conversations. We evaluate the full attack pipeline: whether poisoned memories are written, later retrieved, and ultimately used to steer the following conversations. Across stateful LLM assistants, poisoned memories were added up to 99.8% on GPT-5.5 and 95% on Kimi-K2.6. Crucially, among successful retrievals, poisoned memories cause attacker-intended agentic actions in 60-89% of evaluations across models. These results show that persistent memory can act as a long-term attack surface across multiple future conversations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_15338 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Hidden in Memory: Sleeper Memory Poisoning in LLM Agents Pulipaka, Sidharth Hlebik, Stanislau Raghav, Leonidas Abdelnabi, Sahar Raina, Vyas Sheth, Ivaxi Fritz, Mario Cryptography and Security Artificial Intelligence D.4.6; I.2.7; I.2.11 Large language models are increasingly augmented with persistent memory, allowing assistants to store user-specific information across sessions for personalization and continuity. This statefulness introduces a new security risk: adversarial content can corrupt what an assistant remembers and thereby influence future interactions. We propose and study sleeper memory poisoning, a delayed attack in which an adversary manipulates external context, such as a document, webpage, or repository, to cause the assistant to store a fabricated memory about the user. Unlike conventional prompt injection, the attack can remain dormant and re-emerge across multiple later conversations. We evaluate the full attack pipeline: whether poisoned memories are written, later retrieved, and ultimately used to steer the following conversations. Across stateful LLM assistants, poisoned memories were added up to 99.8% on GPT-5.5 and 95% on Kimi-K2.6. Crucially, among successful retrievals, poisoned memories cause attacker-intended agentic actions in 60-89% of evaluations across models. These results show that persistent memory can act as a long-term attack surface across multiple future conversations. |
| title | Hidden in Memory: Sleeper Memory Poisoning in LLM Agents |
| topic | Cryptography and Security Artificial Intelligence D.4.6; I.2.7; I.2.11 |
| url | https://arxiv.org/abs/2605.15338 |