Dissecting Adversarial Robustness of Multimodal LM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Chen Henry, Shah, Rishi, Koh, Jing Yu, Salakhutdinov, Ruslan, Fried, Daniel, Raghunathan, Aditi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929698540355584
author Wu, Chen Henry
Shah, Rishi
Koh, Jing Yu
Salakhutdinov, Ruslan
Fried, Daniel
Raghunathan, Aditi
author_facet Wu, Chen Henry
Shah, Rishi
Koh, Jing Yu
Salakhutdinov, Ruslan
Fried, Daniel
Raghunathan, Aditi
contents As language models (LMs) are used to build autonomous agents in real environments, ensuring their adversarial robustness becomes a critical challenge. Unlike chatbots, agents are compound systems with multiple components taking actions, which existing LMs safety evaluations do not adequately address. To bridge this gap, we manually create 200 targeted adversarial tasks and evaluation scripts in a realistic threat model on top of VisualWebArena, a real environment for web agents. To systematically examine the robustness of agents, we propose the Agent Robustness Evaluation (ARE) framework. ARE views the agent as a graph showing the flow of intermediate outputs between components and decomposes robustness as the flow of adversarial information on the graph. We find that we can successfully break latest agents that use black-box frontier LMs, including those that perform reflection and tree search. With imperceptible perturbations to a single image (less than 5% of total web page pixels), an attacker can hijack these agents to execute targeted adversarial goals with success rates up to 67%. We also use ARE to rigorously evaluate how the robustness changes as new components are added. We find that inference-time compute that typically improves benign performance can open up new vulnerabilities and harm robustness. An attacker can compromise the evaluator used by the reflexion agent and the value function of the tree search agent, which increases the attack success relatively by 15% and 20%. Our data and code for attacks, defenses, and evaluation are at https://github.com/ChenWu98/agent-attack
format Preprint
id arxiv_https___arxiv_org_abs_2406_12814
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dissecting Adversarial Robustness of Multimodal LM Agents
Wu, Chen Henry
Shah, Rishi
Koh, Jing Yu
Salakhutdinov, Ruslan
Fried, Daniel
Raghunathan, Aditi
Machine Learning
Computation and Language
Cryptography and Security
Computer Vision and Pattern Recognition
As language models (LMs) are used to build autonomous agents in real environments, ensuring their adversarial robustness becomes a critical challenge. Unlike chatbots, agents are compound systems with multiple components taking actions, which existing LMs safety evaluations do not adequately address. To bridge this gap, we manually create 200 targeted adversarial tasks and evaluation scripts in a realistic threat model on top of VisualWebArena, a real environment for web agents. To systematically examine the robustness of agents, we propose the Agent Robustness Evaluation (ARE) framework. ARE views the agent as a graph showing the flow of intermediate outputs between components and decomposes robustness as the flow of adversarial information on the graph. We find that we can successfully break latest agents that use black-box frontier LMs, including those that perform reflection and tree search. With imperceptible perturbations to a single image (less than 5% of total web page pixels), an attacker can hijack these agents to execute targeted adversarial goals with success rates up to 67%. We also use ARE to rigorously evaluate how the robustness changes as new components are added. We find that inference-time compute that typically improves benign performance can open up new vulnerabilities and harm robustness. An attacker can compromise the evaluator used by the reflexion agent and the value function of the tree search agent, which increases the attack success relatively by 15% and 20%. Our data and code for attacks, defenses, and evaluation are at https://github.com/ChenWu98/agent-attack
title Dissecting Adversarial Robustness of Multimodal LM Agents
topic Machine Learning
Computation and Language
Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.12814