R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912059158953984 |
|---|---|
| author | Yuan, Tongxin He, Zhiwei Dong, Lingzhong Wang, Yiming Zhao, Ruijie Xia, Tian Xu, Lizhen Zhou, Binglin Li, Fangqi Zhang, Zhuosheng Wang, Rui Liu, Gongshen |
| author_facet | Yuan, Tongxin He, Zhiwei Dong, Lingzhong Wang, Yiming Zhao, Ruijie Xia, Tian Xu, Lizhen Zhou, Binglin Li, Fangqi Zhang, Zhuosheng Wang, Rui Liu, Gongshen |
| contents | Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering on the harmlessness of LLM-generated content in most prior studies, this work addresses the imperative need for benchmarking the behavioral safety of LLM agents within diverse environments. We introduce R-Judge, a benchmark crafted to evaluate the proficiency of LLMs in judging and identifying safety risks given agent interaction records. R-Judge comprises 569 records of multi-turn agent interaction, encompassing 27 key risk scenarios among 5 application categories and 10 risk types. It is of high-quality curation with annotated safety labels and risk descriptions. Evaluation of 11 LLMs on R-Judge shows considerable room for enhancing the risk awareness of LLMs: The best-performing model, GPT-4o, achieves 74.42% while no other models significantly exceed the random. Moreover, we reveal that risk awareness in open agent scenarios is a multi-dimensional capability involving knowledge and reasoning, thus challenging for LLMs. With further experiments, we find that fine-tuning on safety judgment significantly improve model performance while straightforward prompting mechanisms fail. R-Judge is publicly available at https://github.com/Lordog/R-Judge. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_10019 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | R-Judge: Benchmarking Safety Risk Awareness for LLM Agents Yuan, Tongxin He, Zhiwei Dong, Lingzhong Wang, Yiming Zhao, Ruijie Xia, Tian Xu, Lizhen Zhou, Binglin Li, Fangqi Zhang, Zhuosheng Wang, Rui Liu, Gongshen Computation and Language Artificial Intelligence Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering on the harmlessness of LLM-generated content in most prior studies, this work addresses the imperative need for benchmarking the behavioral safety of LLM agents within diverse environments. We introduce R-Judge, a benchmark crafted to evaluate the proficiency of LLMs in judging and identifying safety risks given agent interaction records. R-Judge comprises 569 records of multi-turn agent interaction, encompassing 27 key risk scenarios among 5 application categories and 10 risk types. It is of high-quality curation with annotated safety labels and risk descriptions. Evaluation of 11 LLMs on R-Judge shows considerable room for enhancing the risk awareness of LLMs: The best-performing model, GPT-4o, achieves 74.42% while no other models significantly exceed the random. Moreover, we reveal that risk awareness in open agent scenarios is a multi-dimensional capability involving knowledge and reasoning, thus challenging for LLMs. With further experiments, we find that fine-tuning on safety judgment significantly improve model performance while straightforward prompting mechanisms fail. R-Judge is publicly available at https://github.com/Lordog/R-Judge. |
| title | R-Judge: Benchmarking Safety Risk Awareness for LLM Agents |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2401.10019 |