AdvAgent: Controllable Blackbox Red-teaming on Web Agents

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Chejian, Kang, Mintong, Zhang, Jiawei, Liao, Zeyi, Mo, Lingbo, Yuan, Mengqi, Sun, Huan, Li, Bo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908387081453568
author Xu, Chejian
Kang, Mintong
Zhang, Jiawei
Liao, Zeyi
Mo, Lingbo
Yuan, Mengqi
Sun, Huan
Li, Bo
author_facet Xu, Chejian
Kang, Mintong
Zhang, Jiawei
Liao, Zeyi
Mo, Lingbo
Yuan, Mengqi
Sun, Huan
Li, Bo
contents Foundation model-based agents are increasingly used to automate complex tasks, enhancing efficiency and productivity. However, their access to sensitive resources and autonomous decision-making also introduce significant security risks, where successful attacks could lead to severe consequences. To systematically uncover these vulnerabilities, we propose AdvAgent, a black-box red-teaming framework for attacking web agents. Unlike existing approaches, AdvAgent employs a reinforcement learning-based pipeline to train an adversarial prompter model that optimizes adversarial prompts using feedback from the black-box agent. With careful attack design, these prompts effectively exploit agent weaknesses while maintaining stealthiness and controllability. Extensive evaluations demonstrate that AdvAgent achieves high success rates against state-of-the-art GPT-4-based web agents across diverse web tasks. Furthermore, we find that existing prompt-based defenses provide only limited protection, leaving agents vulnerable to our framework. These findings highlight critical vulnerabilities in current web agents and emphasize the urgent need for stronger defense mechanisms. We release code at https://ai-secure.github.io/AdvAgent/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17401
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AdvAgent: Controllable Blackbox Red-teaming on Web Agents
Xu, Chejian
Kang, Mintong
Zhang, Jiawei
Liao, Zeyi
Mo, Lingbo
Yuan, Mengqi
Sun, Huan
Li, Bo
Cryptography and Security
Computation and Language
Foundation model-based agents are increasingly used to automate complex tasks, enhancing efficiency and productivity. However, their access to sensitive resources and autonomous decision-making also introduce significant security risks, where successful attacks could lead to severe consequences. To systematically uncover these vulnerabilities, we propose AdvAgent, a black-box red-teaming framework for attacking web agents. Unlike existing approaches, AdvAgent employs a reinforcement learning-based pipeline to train an adversarial prompter model that optimizes adversarial prompts using feedback from the black-box agent. With careful attack design, these prompts effectively exploit agent weaknesses while maintaining stealthiness and controllability. Extensive evaluations demonstrate that AdvAgent achieves high success rates against state-of-the-art GPT-4-based web agents across diverse web tasks. Furthermore, we find that existing prompt-based defenses provide only limited protection, leaving agents vulnerable to our framework. These findings highlight critical vulnerabilities in current web agents and emphasize the urgent need for stronger defense mechanisms. We release code at https://ai-secure.github.io/AdvAgent/.
title AdvAgent: Controllable Blackbox Red-teaming on Web Agents
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2410.17401