How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dziemian, Mateusz, Lin, Maxwell, Fu, Xiaohan, Nowak, Micha, Winter, Nick, Jones, Eliot, Zou, Andy, Ahmad, Lama, Chaudhuri, Kamalika, Chennabasappa, Sahana, Davies, Xander, Deason, Lauren, Edelman, Benjamin L., Emek, Tanner, Evtimov, Ivan, Gust, Jim, Hamin, Maia, He, Kat, Krawiecka, Klaudia, Patana, Riccardo, Perry, Neil, Peterson, Troy, Qi, Xiangyu, Rando, Javier, Wang, Zifan, Wang, Zihan, Whitman, Spencer, Winsor, Eric, Zharmagambetov, Arman, Fredrikson, Matt, Kolter, Zico
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912969197092864
author Dziemian, Mateusz
Lin, Maxwell
Fu, Xiaohan
Nowak, Micha
Winter, Nick
Jones, Eliot
Zou, Andy
Ahmad, Lama
Chaudhuri, Kamalika
Chennabasappa, Sahana
Davies, Xander
Deason, Lauren
Edelman, Benjamin L.
Emek, Tanner
Evtimov, Ivan
Gust, Jim
Hamin, Maia
He, Kat
Krawiecka, Klaudia
Patana, Riccardo
Perry, Neil
Peterson, Troy
Qi, Xiangyu
Rando, Javier
Wang, Zifan
Wang, Zihan
Whitman, Spencer
Winsor, Eric
Zharmagambetov, Arman
Fredrikson, Matt
Kolter, Zico
author_facet Dziemian, Mateusz
Lin, Maxwell
Fu, Xiaohan
Nowak, Micha
Winter, Nick
Jones, Eliot
Zou, Andy
Ahmad, Lama
Chaudhuri, Kamalika
Chennabasappa, Sahana
Davies, Xander
Deason, Lauren
Edelman, Benjamin L.
Emek, Tanner
Evtimov, Ivan
Gust, Jim
Hamin, Maia
He, Kat
Krawiecka, Klaudia
Patana, Riccardo
Perry, Neil
Peterson, Troy
Qi, Xiangyu
Rando, Javier
Wang, Zifan
Wang, Zihan
Whitman, Spencer
Winsor, Eric
Zharmagambetov, Arman
Fredrikson, Matt
Kolter, Zico
contents LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure to indirect prompt injection attacks, where adversarial instructions embedded in external content manipulate agent behavior without user awareness. A critical but underexplored dimension of this threat is concealment: since users tend to observe only an agent's final response, an attack can conceal its existence by presenting no clue of compromise in the final user facing response while successfully executing harmful actions. This leaves users unaware of the manipulation and likely to accept harmful outcomes as legitimate. We present findings from a large scale public red teaming competition evaluating this dual objective across three agent settings: tool calling, coding, and computer use. The competition attracted 464 participants who submitted 272000 attack attempts against 13 frontier models, yielding 8648 successful attacks across 41 scenarios. All models proved vulnerable, with attack success rates ranging from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro). We identify universal attack strategies that transfer across 21 of 41 behaviors and multiple model families, suggesting fundamental weaknesses in instruction following architectures. Capability and robustness showed weak correlation, with Gemini 2.5 Pro exhibiting both high capability and high vulnerability. To address benchmark saturation and obsoleteness, we will endeavor to deliver quarterly updates through continued red teaming competitions. We open source the competition environment for use in evaluations, along with 95 successful attacks against Qwen that did not transfer to any closed source model. We share model-specific attack data with respective frontier labs and the full dataset with the UK AISI and US CAISI to support robustness research.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15714
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
Dziemian, Mateusz
Lin, Maxwell
Fu, Xiaohan
Nowak, Micha
Winter, Nick
Jones, Eliot
Zou, Andy
Ahmad, Lama
Chaudhuri, Kamalika
Chennabasappa, Sahana
Davies, Xander
Deason, Lauren
Edelman, Benjamin L.
Emek, Tanner
Evtimov, Ivan
Gust, Jim
Hamin, Maia
He, Kat
Krawiecka, Klaudia
Patana, Riccardo
Perry, Neil
Peterson, Troy
Qi, Xiangyu
Rando, Javier
Wang, Zifan
Wang, Zihan
Whitman, Spencer
Winsor, Eric
Zharmagambetov, Arman
Fredrikson, Matt
Kolter, Zico
Cryptography and Security
Artificial Intelligence
LLM based agents are increasingly deployed in high stakes settings where they process external data sources such as emails, documents, and code repositories. This creates exposure to indirect prompt injection attacks, where adversarial instructions embedded in external content manipulate agent behavior without user awareness. A critical but underexplored dimension of this threat is concealment: since users tend to observe only an agent's final response, an attack can conceal its existence by presenting no clue of compromise in the final user facing response while successfully executing harmful actions. This leaves users unaware of the manipulation and likely to accept harmful outcomes as legitimate. We present findings from a large scale public red teaming competition evaluating this dual objective across three agent settings: tool calling, coding, and computer use. The competition attracted 464 participants who submitted 272000 attack attempts against 13 frontier models, yielding 8648 successful attacks across 41 scenarios. All models proved vulnerable, with attack success rates ranging from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro). We identify universal attack strategies that transfer across 21 of 41 behaviors and multiple model families, suggesting fundamental weaknesses in instruction following architectures. Capability and robustness showed weak correlation, with Gemini 2.5 Pro exhibiting both high capability and high vulnerability. To address benchmark saturation and obsoleteness, we will endeavor to deliver quarterly updates through continued red teaming competitions. We open source the competition environment for use in evaluations, along with 95 successful attacks against Qwen that did not transfer to any closed source model. We share model-specific attack data with respective frontier labs and the full dataset with the UK AISI and US CAISI to support robustness research.
title How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2603.15714