LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abdelnabi, Sahar, Fay, Aideen, Salem, Ahmed, Zverev, Egor, Liao, Kai-Chieh, Liu, Chi-Huang, Kuo, Chun-Chih, Weigend, Jannis, Manlangit, Danyael, Apostolov, Alex, Umair, Haris, Donato, João, Kawakita, Masayuki, Mahboob, Athar, Bach, Tran Huu, Chiang, Tsun-Han, Cho, Myeongjin, Choi, Hajin, Kim, Byeonghyeon, Lee, Hyeonjin, Pannell, Benjamin, McCauley, Conor, Russinovich, Mark, Paverd, Andrew, Cherubin, Giovanni
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912425234661376
author Abdelnabi, Sahar
Fay, Aideen
Salem, Ahmed
Zverev, Egor
Liao, Kai-Chieh
Liu, Chi-Huang
Kuo, Chun-Chih
Weigend, Jannis
Manlangit, Danyael
Apostolov, Alex
Umair, Haris
Donato, João
Kawakita, Masayuki
Mahboob, Athar
Bach, Tran Huu
Chiang, Tsun-Han
Cho, Myeongjin
Choi, Hajin
Kim, Byeonghyeon
Lee, Hyeonjin
Pannell, Benjamin
McCauley, Conor
Russinovich, Mark
Paverd, Andrew
Cherubin, Giovanni
author_facet Abdelnabi, Sahar
Fay, Aideen
Salem, Ahmed
Zverev, Egor
Liao, Kai-Chieh
Liu, Chi-Huang
Kuo, Chun-Chih
Weigend, Jannis
Manlangit, Danyael
Apostolov, Alex
Umair, Haris
Donato, João
Kawakita, Masayuki
Mahboob, Athar
Bach, Tran Huu
Chiang, Tsun-Han
Cho, Myeongjin
Choi, Hajin
Kim, Byeonghyeon
Lee, Hyeonjin
Pannell, Benjamin
McCauley, Conor
Russinovich, Mark
Paverd, Andrew
Cherubin, Giovanni
contents Indirect Prompt Injection attacks exploit the inherent limitation of Large Language Models (LLMs) to distinguish between instructions and data in their inputs. Despite numerous defense proposals, the systematic evaluation against adaptive adversaries remains limited, even when successful attacks can have wide security and privacy implications, and many real-world LLM-based applications remain vulnerable. We present the results of LLMail-Inject, a public challenge simulating a realistic scenario in which participants adaptively attempted to inject malicious instructions into emails in order to trigger unauthorized tool calls in an LLM-based email assistant. The challenge spanned multiple defense strategies, LLM architectures, and retrieval configurations, resulting in a dataset of 208,095 unique attack submissions from 839 participants. We release the challenge code, the full dataset of submissions, and our analysis demonstrating how this data can provide new insights into the instruction-data separation problem. We hope this will serve as a foundation for future research towards practical structural solutions to prompt injection.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09956
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
Abdelnabi, Sahar
Fay, Aideen
Salem, Ahmed
Zverev, Egor
Liao, Kai-Chieh
Liu, Chi-Huang
Kuo, Chun-Chih
Weigend, Jannis
Manlangit, Danyael
Apostolov, Alex
Umair, Haris
Donato, João
Kawakita, Masayuki
Mahboob, Athar
Bach, Tran Huu
Chiang, Tsun-Han
Cho, Myeongjin
Choi, Hajin
Kim, Byeonghyeon
Lee, Hyeonjin
Pannell, Benjamin
McCauley, Conor
Russinovich, Mark
Paverd, Andrew
Cherubin, Giovanni
Cryptography and Security
Artificial Intelligence
Indirect Prompt Injection attacks exploit the inherent limitation of Large Language Models (LLMs) to distinguish between instructions and data in their inputs. Despite numerous defense proposals, the systematic evaluation against adaptive adversaries remains limited, even when successful attacks can have wide security and privacy implications, and many real-world LLM-based applications remain vulnerable. We present the results of LLMail-Inject, a public challenge simulating a realistic scenario in which participants adaptively attempted to inject malicious instructions into emails in order to trigger unauthorized tool calls in an LLM-based email assistant. The challenge spanned multiple defense strategies, LLM architectures, and retrieval configurations, resulting in a dataset of 208,095 unique attack submissions from 839 participants. We release the challenge code, the full dataset of submissions, and our analysis demonstrating how this data can provide new insights into the instruction-data separation problem. We hope this will serve as a foundation for future research towards practical structural solutions to prompt injection.
title LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2506.09956