SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Xiangyang, Tian, Yuan, Li, Chunyi, Zhang, Kaiwei, Sun, Wei, Zhai, Guangtao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913999901163520
author Zhu, Xiangyang
Tian, Yuan
Li, Chunyi
Zhang, Kaiwei
Sun, Wei
Zhai, Guangtao
author_facet Zhu, Xiangyang
Tian, Yuan
Li, Chunyi
Zhang, Kaiwei
Sun, Wei
Zhai, Guangtao
contents The rapid proliferation of large language models (LLMs) has intensified the requirement for reliable safety evaluation to uncover model vulnerabilities. To this end, numerous LLM safety evaluation benchmarks are proposed. However, existing benchmarks generally rely on labor-intensive manual curation, which causes excessive time and resource consumption. They also exhibit significant redundancy and limited difficulty. To alleviate these problems, we introduce SafetyFlow, the first agent-flow system designed to automate the construction of LLM safety benchmarks. SafetyFlow can automatically build a comprehensive safety benchmark in only four days without any human intervention by orchestrating seven specialized agents, significantly reducing time and resource cost. Equipped with versatile tools, the agents of SafetyFlow ensure process and cost controllability while integrating human expertise into the automatic pipeline. The final constructed dataset, SafetyFlowBench, contains 23,446 queries with low redundancy and strong discriminative power. Our contribution includes the first fully automated benchmarking pipeline and a comprehensive safety benchmark. We evaluate the safety of 49 advanced LLMs on our dataset and conduct extensive experiments to validate our efficacy and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15526
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
Zhu, Xiangyang
Tian, Yuan
Li, Chunyi
Zhang, Kaiwei
Sun, Wei
Zhai, Guangtao
Computation and Language
The rapid proliferation of large language models (LLMs) has intensified the requirement for reliable safety evaluation to uncover model vulnerabilities. To this end, numerous LLM safety evaluation benchmarks are proposed. However, existing benchmarks generally rely on labor-intensive manual curation, which causes excessive time and resource consumption. They also exhibit significant redundancy and limited difficulty. To alleviate these problems, we introduce SafetyFlow, the first agent-flow system designed to automate the construction of LLM safety benchmarks. SafetyFlow can automatically build a comprehensive safety benchmark in only four days without any human intervention by orchestrating seven specialized agents, significantly reducing time and resource cost. Equipped with versatile tools, the agents of SafetyFlow ensure process and cost controllability while integrating human expertise into the automatic pipeline. The final constructed dataset, SafetyFlowBench, contains 23,446 queries with low redundancy and strong discriminative power. Our contribution includes the first fully automated benchmarking pipeline and a comprehensive safety benchmark. We evaluate the safety of 49 advanced LLMs on our dataset and conduct extensive experiments to validate our efficacy and efficiency.
title SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
topic Computation and Language
url https://arxiv.org/abs/2508.15526