AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Andriushchenko, Maksym, Souly, Alexandra, Dziemian, Mateusz, Duenas, Derek, Lin, Maxwell, Wang, Justin, Hendrycks, Dan, Zou, Andy, Kolter, Zico, Fredrikson, Matt, Winsor, Eric, Wynne, Jerome, Gal, Yarin, Davies, Xander
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915249256398848
author Andriushchenko, Maksym
Souly, Alexandra
Dziemian, Mateusz
Duenas, Derek
Lin, Maxwell
Wang, Justin
Hendrycks, Dan
Zou, Andy
Kolter, Zico
Fredrikson, Matt
Winsor, Eric
Wynne, Jerome
Gal, Yarin
Davies, Xander
author_facet Andriushchenko, Maksym
Souly, Alexandra
Dziemian, Mateusz
Duenas, Derek
Lin, Maxwell
Wang, Justin
Hendrycks, Dan
Zou, Andy
Kolter, Zico
Fredrikson, Matt
Winsor, Eric
Wynne, Jerome
Gal, Yarin
Davies, Xander
contents The robustness of LLMs to jailbreak attacks, where users design prompts to circumvent safety measures and misuse model capabilities, has been studied primarily for LLMs acting as simple chatbots. Meanwhile, LLM agents -- which use external tools and can execute multi-stage tasks -- may pose a greater risk if misused, but their robustness remains underexplored. To facilitate research on LLM agent misuse, we propose a new benchmark called AgentHarm. The benchmark includes a diverse set of 110 explicitly malicious agent tasks (440 with augmentations), covering 11 harm categories including fraud, cybercrime, and harassment. In addition to measuring whether models refuse harmful agentic requests, scoring well on AgentHarm requires jailbroken agents to maintain their capabilities following an attack to complete a multi-step task. We evaluate a range of leading LLMs, and find (1) leading LLMs are surprisingly compliant with malicious agent requests without jailbreaking, (2) simple universal jailbreak templates can be adapted to effectively jailbreak agents, and (3) these jailbreaks enable coherent and malicious multi-step agent behavior and retain model capabilities. To enable simple and reliable evaluation of attacks and defenses for LLM-based agents, we publicly release AgentHarm at https://huggingface.co/datasets/ai-safety-institute/AgentHarm.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09024
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Andriushchenko, Maksym
Souly, Alexandra
Dziemian, Mateusz
Duenas, Derek
Lin, Maxwell
Wang, Justin
Hendrycks, Dan
Zou, Andy
Kolter, Zico
Fredrikson, Matt
Winsor, Eric
Wynne, Jerome
Gal, Yarin
Davies, Xander
Machine Learning
Artificial Intelligence
Computation and Language
The robustness of LLMs to jailbreak attacks, where users design prompts to circumvent safety measures and misuse model capabilities, has been studied primarily for LLMs acting as simple chatbots. Meanwhile, LLM agents -- which use external tools and can execute multi-stage tasks -- may pose a greater risk if misused, but their robustness remains underexplored. To facilitate research on LLM agent misuse, we propose a new benchmark called AgentHarm. The benchmark includes a diverse set of 110 explicitly malicious agent tasks (440 with augmentations), covering 11 harm categories including fraud, cybercrime, and harassment. In addition to measuring whether models refuse harmful agentic requests, scoring well on AgentHarm requires jailbroken agents to maintain their capabilities following an attack to complete a multi-step task. We evaluate a range of leading LLMs, and find (1) leading LLMs are surprisingly compliant with malicious agent requests without jailbreaking, (2) simple universal jailbreak templates can be adapted to effectively jailbreak agents, and (3) these jailbreaks enable coherent and malicious multi-step agent behavior and retain model capabilities. To enable simple and reliable evaluation of attacks and defenses for LLM-based agents, we publicly release AgentHarm at https://huggingface.co/datasets/ai-safety-institute/AgentHarm.
title AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.09024