A Critical Evaluation of Defenses against Prompt Injection Attacks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jia, Yuqi, Shao, Zedian, Liu, Yupei, Jia, Jinyuan, Song, Dawn, Gong, Neil Zhenqiang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916756342177792
author Jia, Yuqi
Shao, Zedian
Liu, Yupei
Jia, Jinyuan
Song, Dawn
Gong, Neil Zhenqiang
author_facet Jia, Yuqi
Shao, Zedian
Liu, Yupei
Jia, Jinyuan
Song, Dawn
Gong, Neil Zhenqiang
contents Large Language Models (LLMs) are vulnerable to prompt injection attacks, and several defenses have recently been proposed, often claiming to mitigate these attacks successfully. However, we argue that existing studies lack a principled approach to evaluating these defenses. In this paper, we argue the need to assess defenses across two critical dimensions: (1) effectiveness, measured against both existing and adaptive prompt injection attacks involving diverse target and injected prompts, and (2) general-purpose utility, ensuring that the defense does not compromise the foundational capabilities of the LLM. Our critical evaluation reveals that prior studies have not followed such a comprehensive evaluation methodology. When assessed using this principled approach, we show that existing defenses are not as successful as previously reported. This work provides a foundation for evaluating future defenses and guiding their development. Our code and data are available at: https://github.com/PIEval123/PIEval.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Critical Evaluation of Defenses against Prompt Injection Attacks
Jia, Yuqi
Shao, Zedian
Liu, Yupei
Jia, Jinyuan
Song, Dawn
Gong, Neil Zhenqiang
Cryptography and Security
Artificial Intelligence
Large Language Models (LLMs) are vulnerable to prompt injection attacks, and several defenses have recently been proposed, often claiming to mitigate these attacks successfully. However, we argue that existing studies lack a principled approach to evaluating these defenses. In this paper, we argue the need to assess defenses across two critical dimensions: (1) effectiveness, measured against both existing and adaptive prompt injection attacks involving diverse target and injected prompts, and (2) general-purpose utility, ensuring that the defense does not compromise the foundational capabilities of the LLM. Our critical evaluation reveals that prior studies have not followed such a comprehensive evaluation methodology. When assessed using this principled approach, we show that existing defenses are not as successful as previously reported. This work provides a foundation for evaluating future defenses and guiding their development. Our code and data are available at: https://github.com/PIEval123/PIEval.
title A Critical Evaluation of Defenses against Prompt Injection Attacks
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2505.18333