How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912969197092864 |
|---|---|
| author | Dziemian, Mateusz Lin, Maxwell Fu, Xiaohan Nowak, Micha Winter, Nick Jones, Eliot Zou, Andy Ahmad, Lama Chaudhuri, Kamalika Chennabasappa, Sahana Davies, Xander Deason, Lauren Edelman, Benjamin L. Emek, Tanner Evtimov, Ivan Gust, Jim Hamin, Maia He, Kat Krawiecka, Klaudia Patana, Riccardo Perry, Neil Peterson, Troy Qi, Xiangyu Rando, Javier Wang, Zifan Wang, Zihan Whitman, Spencer Winsor, Eric Zharmagambetov, Arman Fredrikson, Matt Kolter, Zico |
| author_facet | Dziemian, Mateusz Lin, Maxwell Fu, Xiaohan Nowak, Micha Winter, Nick Jones, Eliot Zou, Andy Ahmad, Lama Chaudhuri, Kamalika Chennabasappa, Sahana Davies, Xander Deason, Lauren Edelman, Benjamin L. Emek, Tanner Evtimov, Ivan Gust, Jim Hamin, Maia He, Kat Krawiecka, Klaudia Patana, Riccardo Perry, Neil Peterson, Troy Qi, Xiangyu Rando, Javier Wang, Zifan Wang, Zihan Whitman, Spencer Winsor, Eric Zharmagambetov, Arman Fredrikson, Matt Kolter, Zico |
| contents | LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure to indirect prompt injection attacks, where adversarial instructions embedded in external content manipulate agent behavior without user awareness. A critical but underexplored dimension of this threat is concealment: since users tend to observe only an agent's final response, an attack can conceal its existence by presenting no clue of compromise in the final user facing response while successfully executing harmful actions. This leaves users unaware of the manipulation and likely to accept harmful outcomes as legitimate. We present findings from a large scale public red teaming competition evaluating this dual objective across three agent settings: tool calling, coding, and computer use. The competition attracted 464 participants who submitted 272000 attack attempts against 13 frontier models, yielding 8648 successful attacks across 41 scenarios. All models proved vulnerable, with attack success rates ranging from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro). We identify universal attack strategies that transfer across 21 of 41 behaviors and multiple model families, suggesting fundamental weaknesses in instruction following architectures. Capability and robustness showed weak correlation, with Gemini 2.5 Pro exhibiting both high capability and high vulnerability. To address benchmark saturation and obsoleteness, we will endeavor to deliver quarterly updates through continued red teaming competitions. We open source the competition environment for use in evaluations, along with 95 successful attacks against Qwen that did not transfer to any closed source model. We share model-specific attack data with respective frontier labs and the full dataset with the UK AISI and US CAISI to support robustness research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_15714 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition Dziemian, Mateusz Lin, Maxwell Fu, Xiaohan Nowak, Micha Winter, Nick Jones, Eliot Zou, Andy Ahmad, Lama Chaudhuri, Kamalika Chennabasappa, Sahana Davies, Xander Deason, Lauren Edelman, Benjamin L. Emek, Tanner Evtimov, Ivan Gust, Jim Hamin, Maia He, Kat Krawiecka, Klaudia Patana, Riccardo Perry, Neil Peterson, Troy Qi, Xiangyu Rando, Javier Wang, Zifan Wang, Zihan Whitman, Spencer Winsor, Eric Zharmagambetov, Arman Fredrikson, Matt Kolter, Zico Cryptography and Security Artificial Intelligence LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure to indirect prompt injection attacks, where adversarial instructions embedded in external content manipulate agent behavior without user awareness. A critical but underexplored dimension of this threat is concealment: since users tend to observe only an agent's final response, an attack can conceal its existence by presenting no clue of compromise in the final user facing response while successfully executing harmful actions. This leaves users unaware of the manipulation and likely to accept harmful outcomes as legitimate. We present findings from a large scale public red teaming competition evaluating this dual objective across three agent settings: tool calling, coding, and computer use. The competition attracted 464 participants who submitted 272000 attack attempts against 13 frontier models, yielding 8648 successful attacks across 41 scenarios. All models proved vulnerable, with attack success rates ranging from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro). We identify universal attack strategies that transfer across 21 of 41 behaviors and multiple model families, suggesting fundamental weaknesses in instruction following architectures. Capability and robustness showed weak correlation, with Gemini 2.5 Pro exhibiting both high capability and high vulnerability. To address benchmark saturation and obsoleteness, we will endeavor to deliver quarterly updates through continued red teaming competitions. We open source the competition environment for use in evaluations, along with 95 successful attacks against Qwen that did not transfer to any closed source model. We share model-specific attack data with respective frontier labs and the full dataset with the UK AISI and US CAISI to support robustness research. |
| title | How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2603.15714 |