Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yi, Jingwei, Xie, Yueqi, Zhu, Bin, Kiciman, Emre, Sun, Guangzhong, Xie, Xing, Wu, Fangzhao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912204357369856
author Yi, Jingwei
Xie, Yueqi
Zhu, Bin
Kiciman, Emre
Sun, Guangzhong
Xie, Xing
Wu, Fangzhao
author_facet Yi, Jingwei
Xie, Yueqi
Zhu, Bin
Kiciman, Emre
Sun, Guangzhong
Xie, Xing
Wu, Fangzhao
contents The integration of large language models with external content has enabled applications such as Microsoft Copilot but also introduced vulnerabilities to indirect prompt injection attacks. In these attacks, malicious instructions embedded within external content can manipulate LLM outputs, causing deviations from user expectations. To address this critical yet under-explored issue, we introduce the first benchmark for indirect prompt injection attacks, named BIPIA, to assess the risk of such vulnerabilities. Using BIPIA, we evaluate existing LLMs and find them universally vulnerable. Our analysis identifies two key factors contributing to their success: LLMs' inability to distinguish between informational context and actionable instructions, and their lack of awareness in avoiding the execution of instructions within external content. Based on these findings, we propose two novel defense mechanisms-boundary awareness and explicit reminder-to address these vulnerabilities in both black-box and white-box settings. Extensive experiments demonstrate that our black-box defense provides substantial mitigation, while our white-box defense reduces the attack success rate to near-zero levels, all while preserving the output quality of LLMs. We hope this work inspires further research into securing LLM applications and fostering their safe and reliable use.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14197
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
Yi, Jingwei
Xie, Yueqi
Zhu, Bin
Kiciman, Emre
Sun, Guangzhong
Xie, Xing
Wu, Fangzhao
Computation and Language
Artificial Intelligence
The integration of large language models with external content has enabled applications such as Microsoft Copilot but also introduced vulnerabilities to indirect prompt injection attacks. In these attacks, malicious instructions embedded within external content can manipulate LLM outputs, causing deviations from user expectations. To address this critical yet under-explored issue, we introduce the first benchmark for indirect prompt injection attacks, named BIPIA, to assess the risk of such vulnerabilities. Using BIPIA, we evaluate existing LLMs and find them universally vulnerable. Our analysis identifies two key factors contributing to their success: LLMs' inability to distinguish between informational context and actionable instructions, and their lack of awareness in avoiding the execution of instructions within external content. Based on these findings, we propose two novel defense mechanisms-boundary awareness and explicit reminder-to address these vulnerabilities in both black-box and white-box settings. Extensive experiments demonstrate that our black-box defense provides substantial mitigation, while our white-box defense reduces the attack success rate to near-zero levels, all while preserving the output quality of LLMs. We hope this work inspires further research into securing LLM applications and fostering their safe and reliable use.
title Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2312.14197