Context manipulation attacks : Web agents are susceptible to corrupted memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patlan, Atharv Singh, Hebbar, Ashwin, Viswanath, Pramod, Mittal, Prateek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909654505750528
author Patlan, Atharv Singh
Hebbar, Ashwin
Viswanath, Pramod
Mittal, Prateek
author_facet Patlan, Atharv Singh
Hebbar, Ashwin
Viswanath, Pramod
Mittal, Prateek
contents Autonomous web navigation agents, which translate natural language instructions into sequences of browser actions, are increasingly deployed for complex tasks across e-commerce, information retrieval, and content discovery. Due to the stateless nature of large language models (LLMs), these agents rely heavily on external memory systems to maintain context across interactions. Unlike centralized systems where context is securely stored server-side, agent memory is often managed client-side or by third-party applications, creating significant security vulnerabilities. This was recently exploited to attack production systems. We introduce and formalize "plan injection," a novel context manipulation attack that corrupts these agents' internal task representations by targeting this vulnerable context. Through systematic evaluation of two popular web agents, Browser-use and Agent-E, we show that plan injections bypass robust prompt injection defenses, achieving up to 3x higher attack success rates than comparable prompt-based attacks. Furthermore, "context-chained injections," which craft logical bridges between legitimate user goals and attacker objectives, lead to a 17.7% increase in success rate for privacy exfiltration tasks. Our findings highlight that secure memory handling must be a first-class concern in agentic systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17318
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context manipulation attacks : Web agents are susceptible to corrupted memory
Patlan, Atharv Singh
Hebbar, Ashwin
Viswanath, Pramod
Mittal, Prateek
Cryptography and Security
Artificial Intelligence
Autonomous web navigation agents, which translate natural language instructions into sequences of browser actions, are increasingly deployed for complex tasks across e-commerce, information retrieval, and content discovery. Due to the stateless nature of large language models (LLMs), these agents rely heavily on external memory systems to maintain context across interactions. Unlike centralized systems where context is securely stored server-side, agent memory is often managed client-side or by third-party applications, creating significant security vulnerabilities. This was recently exploited to attack production systems. We introduce and formalize "plan injection," a novel context manipulation attack that corrupts these agents' internal task representations by targeting this vulnerable context. Through systematic evaluation of two popular web agents, Browser-use and Agent-E, we show that plan injections bypass robust prompt injection defenses, achieving up to 3x higher attack success rates than comparable prompt-based attacks. Furthermore, "context-chained injections," which craft logical bridges between legitimate user goals and attacker objectives, lead to a 17.7% increase in success rate for privacy exfiltration tasks. Our findings highlight that secure memory handling must be a first-class concern in agentic systems.
title Context manipulation attacks : Web agents are susceptible to corrupted memory
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2506.17318