SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Sheng, Pang, Xianghe, Ding, Yuanzhuo, Chen, Menglan, Bi, Yutong, Xiong, Yichen, Huang, Wenhao, Xiang, Zhen, Shao, Jing, Chen, Siheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908621572407296
author Yin, Sheng
Pang, Xianghe
Ding, Yuanzhuo
Chen, Menglan
Bi, Yutong
Xiong, Yichen
Huang, Wenhao
Xiang, Zhen
Shao, Jing
Chen, Siheng
author_facet Yin, Sheng
Pang, Xianghe
Ding, Yuanzhuo
Chen, Menglan
Bi, Yutong
Xiong, Yichen
Huang, Wenhao
Xiang, Zhen
Shao, Jing
Chen, Siheng
contents With the integration of large language models (LLMs), embodied agents have strong capabilities to understand and plan complicated natural language instructions. However, a foreseeable issue is that those embodied agents can also flawlessly execute some hazardous tasks, potentially causing damages in the real world. Existing benchmarks predominantly overlook critical safety risks, focusing solely on planning performance, while a few evaluate LLMs' safety awareness only on non-interactive image-text data. To address this gap, we present SafeAgentBench -- the first comprehensive benchmark for safety-aware task planning of embodied LLM agents in interactive simulation environments, covering both explicit and implicit hazards. SafeAgentBench includes: (1) an executable, diverse, and high-quality dataset of 750 tasks, rigorously curated to cover 10 potential hazards and 3 task types; (2) SafeAgentEnv, a universal embodied environment with a low-level controller, supporting multi-agent execution with 17 high-level actions for 9 state-of-the-art baselines; and (3) reliable evaluation methods from both execution and semantic perspectives. Experimental results show that, although agents based on different design frameworks exhibit substantial differences in task success rates, their overall safety awareness remains weak. The most safety-conscious baseline achieves only a 10% rejection rate for detailed hazardous tasks. Moreover, simply replacing the LLM driving the agent does not lead to notable improvements in safety awareness. Dataset and codes are available in https://github.com/shengyin1224/SafeAgentBench and https://huggingface.co/datasets/safeagentbench/SafeAgentBench.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13178
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
Yin, Sheng
Pang, Xianghe
Ding, Yuanzhuo
Chen, Menglan
Bi, Yutong
Xiong, Yichen
Huang, Wenhao
Xiang, Zhen
Shao, Jing
Chen, Siheng
Cryptography and Security
Artificial Intelligence
Robotics
With the integration of large language models (LLMs), embodied agents have strong capabilities to understand and plan complicated natural language instructions. However, a foreseeable issue is that those embodied agents can also flawlessly execute some hazardous tasks, potentially causing damages in the real world. Existing benchmarks predominantly overlook critical safety risks, focusing solely on planning performance, while a few evaluate LLMs' safety awareness only on non-interactive image-text data. To address this gap, we present SafeAgentBench -- the first comprehensive benchmark for safety-aware task planning of embodied LLM agents in interactive simulation environments, covering both explicit and implicit hazards. SafeAgentBench includes: (1) an executable, diverse, and high-quality dataset of 750 tasks, rigorously curated to cover 10 potential hazards and 3 task types; (2) SafeAgentEnv, a universal embodied environment with a low-level controller, supporting multi-agent execution with 17 high-level actions for 9 state-of-the-art baselines; and (3) reliable evaluation methods from both execution and semantic perspectives. Experimental results show that, although agents based on different design frameworks exhibit substantial differences in task success rates, their overall safety awareness remains weak. The most safety-conscious baseline achieves only a 10% rejection rate for detailed hazardous tasks. Moreover, simply replacing the LLM driving the agent does not lead to notable improvements in safety awareness. Dataset and codes are available in https://github.com/shengyin1224/SafeAgentBench and https://huggingface.co/datasets/safeagentbench/SafeAgentBench.
title SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
topic Cryptography and Security
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2412.13178