How Not to Detect Prompt Injections with an LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choudhary, Sarthak, Anshumaan, Divyam, Palumbo, Nils, Jha, Somesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911305813721088
author Choudhary, Sarthak
Anshumaan, Divyam
Palumbo, Nils
Jha, Somesh
author_facet Choudhary, Sarthak
Anshumaan, Divyam
Palumbo, Nils
Jha, Somesh
contents LLM-integrated applications and agents are vulnerable to prompt injection attacks, where adversaries embed malicious instructions within seemingly benign input data to manipulate the LLM's intended behavior. Recent defenses based on known-answer detection (KAD) scheme have reported near-perfect performance by observing an LLM's output to classify input data as clean or contaminated. KAD attempts to repurpose the very susceptibility to prompt injection as a defensive mechanism. We formally characterize the KAD scheme and uncover a structural vulnerability that invalidates its core security premise. To exploit this fundamental vulnerability, we methodically design an adaptive attack, DataFlip. It consistently evades KAD defenses, achieving detection rates as low as $0\%$ while reliably inducing malicious behavior with a success rate of $91\%$, all without requiring white-box access to the LLM or any optimization procedures.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Not to Detect Prompt Injections with an LLM
Choudhary, Sarthak
Anshumaan, Divyam
Palumbo, Nils
Jha, Somesh
Cryptography and Security
Artificial Intelligence
Machine Learning
LLM-integrated applications and agents are vulnerable to prompt injection attacks, where adversaries embed malicious instructions within seemingly benign input data to manipulate the LLM's intended behavior. Recent defenses based on known-answer detection (KAD) scheme have reported near-perfect performance by observing an LLM's output to classify input data as clean or contaminated. KAD attempts to repurpose the very susceptibility to prompt injection as a defensive mechanism. We formally characterize the KAD scheme and uncover a structural vulnerability that invalidates its core security premise. To exploit this fundamental vulnerability, we methodically design an adaptive attack, DataFlip. It consistently evades KAD defenses, achieving detection rates as low as $0\%$ while reliably inducing malicious behavior with a success rate of $91\%$, all without requiring white-box access to the LLM or any optimization procedures.
title How Not to Detect Prompt Injections with an LLM
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.05630