RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xiuyuan, Zhao, Jian, Yuan, Yuchen, Zhang, Tianle, Zhou, Huilin, Zhu, Zheng, Hu, Ping, Kong, Linghe, Zhang, Chi, Huang, Weiran, Li, Xuelong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911227490336768
author Chen, Xiuyuan
Zhao, Jian
Yuan, Yuchen
Zhang, Tianle
Zhou, Huilin
Zhu, Zheng
Hu, Ping
Kong, Linghe
Zhang, Chi
Huang, Weiran
Li, Xuelong
author_facet Chen, Xiuyuan
Zhao, Jian
Yuan, Yuchen
Zhang, Tianle
Zhou, Huilin
Zhu, Zheng
Hu, Ping
Kong, Linghe
Zhang, Chi
Huang, Weiran
Li, Xuelong
contents Existing safety evaluation methods for large language models (LLMs) suffer from inherent limitations, including evaluator bias and detection failures arising from model homogeneity, which collectively undermine the robustness of risk evaluation processes. This paper seeks to re-examine the risk evaluation paradigm by introducing a theoretical framework that reconstructs the underlying risk concept space. Specifically, we decompose the latent risk concept space into three mutually exclusive subspaces: the explicit risk subspace (encompassing direct violations of safety guidelines), the implicit risk subspace (capturing potential malicious content that requires contextual reasoning for identification), and the non-risk subspace. Furthermore, we propose RADAR, a multi-agent collaborative evaluation framework that leverages multi-round debate mechanisms through four specialized complementary roles and employs dynamic update mechanisms to achieve self-evolution of risk concept distributions. This approach enables comprehensive coverage of both explicit and implicit risks while mitigating evaluator bias. To validate the effectiveness of our framework, we construct an evaluation dataset comprising 800 challenging cases. Extensive experiments on our challenging testset and public benchmarks demonstrate that RADAR significantly outperforms baseline evaluation methods across multiple dimensions, including accuracy, stability, and self-evaluation risk sensitivity. Notably, RADAR achieves a 28.87% improvement in risk identification accuracy compared to the strongest baseline evaluation method.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25271
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
Chen, Xiuyuan
Zhao, Jian
Yuan, Yuchen
Zhang, Tianle
Zhou, Huilin
Zhu, Zheng
Hu, Ping
Kong, Linghe
Zhang, Chi
Huang, Weiran
Li, Xuelong
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multiagent Systems
Existing safety evaluation methods for large language models (LLMs) suffer from inherent limitations, including evaluator bias and detection failures arising from model homogeneity, which collectively undermine the robustness of risk evaluation processes. This paper seeks to re-examine the risk evaluation paradigm by introducing a theoretical framework that reconstructs the underlying risk concept space. Specifically, we decompose the latent risk concept space into three mutually exclusive subspaces: the explicit risk subspace (encompassing direct violations of safety guidelines), the implicit risk subspace (capturing potential malicious content that requires contextual reasoning for identification), and the non-risk subspace. Furthermore, we propose RADAR, a multi-agent collaborative evaluation framework that leverages multi-round debate mechanisms through four specialized complementary roles and employs dynamic update mechanisms to achieve self-evolution of risk concept distributions. This approach enables comprehensive coverage of both explicit and implicit risks while mitigating evaluator bias. To validate the effectiveness of our framework, we construct an evaluation dataset comprising 800 challenging cases. Extensive experiments on our challenging testset and public benchmarks demonstrate that RADAR significantly outperforms baseline evaluation methods across multiple dimensions, including accuracy, stability, and self-evaluation risk sensitivity. Notably, RADAR achieves a 28.87% improvement in risk identification accuracy compared to the strongest baseline evaluation method.
title RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2509.25271