R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Tongxin, He, Zhiwei, Dong, Lingzhong, Wang, Yiming, Zhao, Ruijie, Xia, Tian, Xu, Lizhen, Zhou, Binglin, Li, Fangqi, Zhang, Zhuosheng, Wang, Rui, Liu, Gongshen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912059158953984
author Yuan, Tongxin
He, Zhiwei
Dong, Lingzhong
Wang, Yiming
Zhao, Ruijie
Xia, Tian
Xu, Lizhen
Zhou, Binglin
Li, Fangqi
Zhang, Zhuosheng
Wang, Rui
Liu, Gongshen
author_facet Yuan, Tongxin
He, Zhiwei
Dong, Lingzhong
Wang, Yiming
Zhao, Ruijie
Xia, Tian
Xu, Lizhen
Zhou, Binglin
Li, Fangqi
Zhang, Zhuosheng
Wang, Rui
Liu, Gongshen
contents Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering on the harmlessness of LLM-generated content in most prior studies, this work addresses the imperative need for benchmarking the behavioral safety of LLM agents within diverse environments. We introduce R-Judge, a benchmark crafted to evaluate the proficiency of LLMs in judging and identifying safety risks given agent interaction records. R-Judge comprises 569 records of multi-turn agent interaction, encompassing 27 key risk scenarios among 5 application categories and 10 risk types. It is of high-quality curation with annotated safety labels and risk descriptions. Evaluation of 11 LLMs on R-Judge shows considerable room for enhancing the risk awareness of LLMs: The best-performing model, GPT-4o, achieves 74.42% while no other models significantly exceed the random. Moreover, we reveal that risk awareness in open agent scenarios is a multi-dimensional capability involving knowledge and reasoning, thus challenging for LLMs. With further experiments, we find that fine-tuning on safety judgment significantly improve model performance while straightforward prompting mechanisms fail. R-Judge is publicly available at https://github.com/Lordog/R-Judge.
format Preprint
id arxiv_https___arxiv_org_abs_2401_10019
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
Yuan, Tongxin
He, Zhiwei
Dong, Lingzhong
Wang, Yiming
Zhao, Ruijie
Xia, Tian
Xu, Lizhen
Zhou, Binglin
Li, Fangqi
Zhang, Zhuosheng
Wang, Rui
Liu, Gongshen
Computation and Language
Artificial Intelligence
Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering on the harmlessness of LLM-generated content in most prior studies, this work addresses the imperative need for benchmarking the behavioral safety of LLM agents within diverse environments. We introduce R-Judge, a benchmark crafted to evaluate the proficiency of LLMs in judging and identifying safety risks given agent interaction records. R-Judge comprises 569 records of multi-turn agent interaction, encompassing 27 key risk scenarios among 5 application categories and 10 risk types. It is of high-quality curation with annotated safety labels and risk descriptions. Evaluation of 11 LLMs on R-Judge shows considerable room for enhancing the risk awareness of LLMs: The best-performing model, GPT-4o, achieves 74.42% while no other models significantly exceed the random. Moreover, we reveal that risk awareness in open agent scenarios is a multi-dimensional capability involving knowledge and reasoning, thus challenging for LLMs. With further experiments, we find that fine-tuning on safety judgment significantly improve model performance while straightforward prompting mechanisms fail. R-Judge is publicly available at https://github.com/Lordog/R-Judge.
title R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.10019