WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yinuo, Xu, Ruohan, Wang, Xilong, Jia, Yuqi, Gong, Neil Zhenqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909820326510592
author Liu, Yinuo
Xu, Ruohan
Wang, Xilong
Jia, Yuqi
Gong, Neil Zhenqiang
author_facet Liu, Yinuo
Xu, Ruohan
Wang, Xilong
Jia, Yuqi
Gong, Neil Zhenqiang
contents Multiple prompt injection attacks have been proposed against web agents. At the same time, various methods have been developed to detect general prompt injection attacks, but none have been systematically evaluated for web agents. In this work, we bridge this gap by presenting the first comprehensive benchmark study on detecting prompt injection attacks targeting web agents. We begin by introducing a fine-grained categorization of such attacks based on the threat model. We then construct datasets containing both malicious and benign samples: malicious text segments generated by different attacks, benign text segments from four categories, malicious images produced by attacks, and benign images from two categories. Next, we systematize both text-based and image-based detection methods. Finally, we evaluate their performance across multiple scenarios. Our key findings show that while some detectors can identify attacks that rely on explicit textual instructions or visible image perturbations with moderate to high accuracy, they largely fail against attacks that omit explicit instructions or employ imperceptible perturbations. Our datasets and code are released at: https://github.com/Norrrrrrr-lyn/WAInjectBench.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
Liu, Yinuo
Xu, Ruohan
Wang, Xilong
Jia, Yuqi
Gong, Neil Zhenqiang
Cryptography and Security
Artificial Intelligence
Computation and Language
Multiple prompt injection attacks have been proposed against web agents. At the same time, various methods have been developed to detect general prompt injection attacks, but none have been systematically evaluated for web agents. In this work, we bridge this gap by presenting the first comprehensive benchmark study on detecting prompt injection attacks targeting web agents. We begin by introducing a fine-grained categorization of such attacks based on the threat model. We then construct datasets containing both malicious and benign samples: malicious text segments generated by different attacks, benign text segments from four categories, malicious images produced by attacks, and benign images from two categories. Next, we systematize both text-based and image-based detection methods. Finally, we evaluate their performance across multiple scenarios. Our key findings show that while some detectors can identify attacks that rely on explicit textual instructions or visible image perturbations with moderate to high accuracy, they largely fail against attacks that omit explicit instructions or employ imperceptible perturbations. Our datasets and code are released at: https://github.com/Norrrrrrr-lyn/WAInjectBench.
title WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.01354