AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Feng, Yunhao, Ding, Yifan, Tan, Yingshui, Ma, Xingjun, Li, Yige, Wu, Yutao, Gao, Yifeng, Zhai, Kun, Guo, Yanming
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917382812860416
author Feng, Yunhao
Ding, Yifan
Tan, Yingshui
Ma, Xingjun
Li, Yige
Wu, Yutao
Gao, Yifeng
Zhai, Kun
Guo, Yanming
author_facet Feng, Yunhao
Ding, Yifan
Tan, Yingshui
Ma, Xingjun
Li, Yige
Wu, Yutao
Gao, Yifeng
Zhai, Kun
Guo, Yanming
contents Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across interactions and translate intermediate outputs into concrete actions. This creates a distinct safety challenge in that harmful behavior may emerge through sequences of individually plausible steps, including intermediate actions that appear locally acceptable but collectively lead to unauthorized actions. We present \textbf{AgentHazard}, a benchmark for evaluating harmful behavior in computer-use agents. AgentHazard contains \textbf{2,653} instances spanning diverse risk categories and attack strategies. Each instance pairs a harmful objective with a sequence of operational steps that are locally legitimate but jointly induce unsafe behavior. The benchmark evaluates whether agents can recognize and interrupt harm arising from accumulated context, repeated tool use, intermediate actions, and dependencies across steps. We evaluate AgentHazard on Claude Code, OpenClaw, and IFlow using mostly open or openly deployable models from the Qwen3, Kimi, GLM, and DeepSeek families. Our experimental results indicate that current systems remain highly vulnerable. In particular, when powered by Qwen3-Coder, Claude Code exhibits an attack success rate of \textbf{73.63\%}, suggesting that model alignment alone does not reliably guarantee the safety of autonomous agents.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02947
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
Feng, Yunhao
Ding, Yifan
Tan, Yingshui
Ma, Xingjun
Li, Yige
Wu, Yutao
Gao, Yifeng
Zhai, Kun
Guo, Yanming
Artificial Intelligence
Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across interactions and translate intermediate outputs into concrete actions. This creates a distinct safety challenge in that harmful behavior may emerge through sequences of individually plausible steps, including intermediate actions that appear locally acceptable but collectively lead to unauthorized actions. We present \textbf{AgentHazard}, a benchmark for evaluating harmful behavior in computer-use agents. AgentHazard contains \textbf{2,653} instances spanning diverse risk categories and attack strategies. Each instance pairs a harmful objective with a sequence of operational steps that are locally legitimate but jointly induce unsafe behavior. The benchmark evaluates whether agents can recognize and interrupt harm arising from accumulated context, repeated tool use, intermediate actions, and dependencies across steps. We evaluate AgentHazard on Claude Code, OpenClaw, and IFlow using mostly open or openly deployable models from the Qwen3, Kimi, GLM, and DeepSeek families. Our experimental results indicate that current systems remain highly vulnerable. In particular, when powered by Qwen3-Coder, Claude Code exhibits an attack success rate of \textbf{73.63\%}, suggesting that model alignment alone does not reliably guarantee the safety of autonomous agents.
title AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2604.02947