LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912425234661376 |
|---|---|
| author | Abdelnabi, Sahar Fay, Aideen Salem, Ahmed Zverev, Egor Liao, Kai-Chieh Liu, Chi-Huang Kuo, Chun-Chih Weigend, Jannis Manlangit, Danyael Apostolov, Alex Umair, Haris Donato, João Kawakita, Masayuki Mahboob, Athar Bach, Tran Huu Chiang, Tsun-Han Cho, Myeongjin Choi, Hajin Kim, Byeonghyeon Lee, Hyeonjin Pannell, Benjamin McCauley, Conor Russinovich, Mark Paverd, Andrew Cherubin, Giovanni |
| author_facet | Abdelnabi, Sahar Fay, Aideen Salem, Ahmed Zverev, Egor Liao, Kai-Chieh Liu, Chi-Huang Kuo, Chun-Chih Weigend, Jannis Manlangit, Danyael Apostolov, Alex Umair, Haris Donato, João Kawakita, Masayuki Mahboob, Athar Bach, Tran Huu Chiang, Tsun-Han Cho, Myeongjin Choi, Hajin Kim, Byeonghyeon Lee, Hyeonjin Pannell, Benjamin McCauley, Conor Russinovich, Mark Paverd, Andrew Cherubin, Giovanni |
| contents | Indirect Prompt Injection attacks exploit the inherent limitation of Large Language Models (LLMs) to distinguish between instructions and data in their inputs. Despite numerous defense proposals, the systematic evaluation against adaptive adversaries remains limited, even when successful attacks can have wide security and privacy implications, and many real-world LLM-based applications remain vulnerable. We present the results of LLMail-Inject, a public challenge simulating a realistic scenario in which participants adaptively attempted to inject malicious instructions into emails in order to trigger unauthorized tool calls in an LLM-based email assistant. The challenge spanned multiple defense strategies, LLM architectures, and retrieval configurations, resulting in a dataset of 208,095 unique attack submissions from 839 participants. We release the challenge code, the full dataset of submissions, and our analysis demonstrating how this data can provide new insights into the instruction-data separation problem. We hope this will serve as a foundation for future research towards practical structural solutions to prompt injection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_09956 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge Abdelnabi, Sahar Fay, Aideen Salem, Ahmed Zverev, Egor Liao, Kai-Chieh Liu, Chi-Huang Kuo, Chun-Chih Weigend, Jannis Manlangit, Danyael Apostolov, Alex Umair, Haris Donato, João Kawakita, Masayuki Mahboob, Athar Bach, Tran Huu Chiang, Tsun-Han Cho, Myeongjin Choi, Hajin Kim, Byeonghyeon Lee, Hyeonjin Pannell, Benjamin McCauley, Conor Russinovich, Mark Paverd, Andrew Cherubin, Giovanni Cryptography and Security Artificial Intelligence Indirect Prompt Injection attacks exploit the inherent limitation of Large Language Models (LLMs) to distinguish between instructions and data in their inputs. Despite numerous defense proposals, the systematic evaluation against adaptive adversaries remains limited, even when successful attacks can have wide security and privacy implications, and many real-world LLM-based applications remain vulnerable. We present the results of LLMail-Inject, a public challenge simulating a realistic scenario in which participants adaptively attempted to inject malicious instructions into emails in order to trigger unauthorized tool calls in an LLM-based email assistant. The challenge spanned multiple defense strategies, LLM architectures, and retrieval configurations, resulting in a dataset of 208,095 unique attack submissions from 839 participants. We release the challenge code, the full dataset of submissions, and our analysis demonstrating how this data can provide new insights into the instruction-data separation problem. We hope this will serve as a foundation for future research towards practical structural solutions to prompt injection. |
| title | LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2506.09956 |