SafeArena: Evaluating the Safety of Autonomous Web Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tur, Ada Defne, Meade, Nicholas, Lù, Xing Han, Zambrano, Alejandra, Patel, Arkil, Durmus, Esin, Gella, Spandana, Stańczak, Karolina, Reddy, Siva
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929745897193472
author Tur, Ada Defne
Meade, Nicholas
Lù, Xing Han
Zambrano, Alejandra
Patel, Arkil
Durmus, Esin
Gella, Spandana
Stańczak, Karolina
Reddy, Siva
author_facet Tur, Ada Defne
Meade, Nicholas
Lù, Xing Han
Zambrano, Alejandra
Patel, Arkil
Durmus, Esin
Gella, Spandana
Stańczak, Karolina
Reddy, Siva
contents LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To evaluate these risks, we propose SafeArena, the first benchmark to focus on the deliberate misuse of web agents. SafeArena comprises 250 safe and 250 harmful tasks across four websites. We classify the harmful tasks into five harm categories -- misinformation, illegal activity, harassment, cybercrime, and social bias, designed to assess realistic misuses of web agents. We evaluate leading LLM-based web agents, including GPT-4o, Claude-3.5 Sonnet, Qwen-2-VL 72B, and Llama-3.2 90B, on our benchmark. To systematically assess their susceptibility to harmful tasks, we introduce the Agent Risk Assessment framework that categorizes agent behavior across four risk levels. We find agents are surprisingly compliant with malicious requests, with GPT-4o and Qwen-2 completing 34.7% and 27.3% of harmful requests, respectively. Our findings highlight the urgent need for safety alignment procedures for web agents. Our benchmark is available here: https://safearena.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2503_04957
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeArena: Evaluating the Safety of Autonomous Web Agents
Tur, Ada Defne
Meade, Nicholas
Lù, Xing Han
Zambrano, Alejandra
Patel, Arkil
Durmus, Esin
Gella, Spandana
Stańczak, Karolina
Reddy, Siva
Machine Learning
Artificial Intelligence
Computation and Language
LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To evaluate these risks, we propose SafeArena, the first benchmark to focus on the deliberate misuse of web agents. SafeArena comprises 250 safe and 250 harmful tasks across four websites. We classify the harmful tasks into five harm categories -- misinformation, illegal activity, harassment, cybercrime, and social bias, designed to assess realistic misuses of web agents. We evaluate leading LLM-based web agents, including GPT-4o, Claude-3.5 Sonnet, Qwen-2-VL 72B, and Llama-3.2 90B, on our benchmark. To systematically assess their susceptibility to harmful tasks, we introduce the Agent Risk Assessment framework that categorizes agent behavior across four risk levels. We find agents are surprisingly compliant with malicious requests, with GPT-4o and Qwen-2 completing 34.7% and 27.3% of harmful requests, respectively. Our findings highlight the urgent need for safety alignment procedures for web agents. Our benchmark is available here: https://safearena.github.io
title SafeArena: Evaluating the Safety of Autonomous Web Agents
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.04957