WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Evtimov, Ivan, Zharmagambetov, Arman, Grattafiori, Aaron, Guo, Chuan, Chaudhuri, Kamalika
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910949808537600
author Evtimov, Ivan
Zharmagambetov, Arman
Grattafiori, Aaron
Guo, Chuan
Chaudhuri, Kamalika
author_facet Evtimov, Ivan
Zharmagambetov, Arman
Grattafiori, Aaron
Guo, Chuan
Chaudhuri, Kamalika
contents Autonomous UI agents powered by AI have tremendous potential to boost human productivity by automating routine tasks such as filing taxes and paying bills. However, a major challenge in unlocking their full potential is security, which is exacerbated by the agent's ability to take action on their user's behalf. Existing tests for prompt injections in web agents either over-simplify the threat by testing unrealistic scenarios or giving the attacker too much power, or look at single-step isolated tasks. To more accurately measure progress for secure web agents, we introduce WASP -- a new publicly available benchmark for end-to-end evaluation of Web Agent Security against Prompt injection attacks. Evaluating with WASP shows that even top-tier AI models, including those with advanced reasoning capabilities, can be deceived by simple, low-effort human-written injections in very realistic scenarios. Our end-to-end evaluation reveals a previously unobserved insight: while attacks partially succeed in up to 86% of the case, even state-of-the-art agents often struggle to fully complete the attacker goals -- highlighting the current state of security by incompetence.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18575
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
Evtimov, Ivan
Zharmagambetov, Arman
Grattafiori, Aaron
Guo, Chuan
Chaudhuri, Kamalika
Cryptography and Security
Artificial Intelligence
Autonomous UI agents powered by AI have tremendous potential to boost human productivity by automating routine tasks such as filing taxes and paying bills. However, a major challenge in unlocking their full potential is security, which is exacerbated by the agent's ability to take action on their user's behalf. Existing tests for prompt injections in web agents either over-simplify the threat by testing unrealistic scenarios or giving the attacker too much power, or look at single-step isolated tasks. To more accurately measure progress for secure web agents, we introduce WASP -- a new publicly available benchmark for end-to-end evaluation of Web Agent Security against Prompt injection attacks. Evaluating with WASP shows that even top-tier AI models, including those with advanced reasoning capabilities, can be deceived by simple, low-effort human-written injections in very realistic scenarios. Our end-to-end evaluation reveals a previously unobserved insight: while attacks partially succeed in up to 86% of the case, even state-of-the-art agents often struggle to fully complete the attacker goals -- highlighting the current state of security by incompetence.
title WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2504.18575