Saved in:
Bibliographic Details
Main Authors: Zhu, Xiangyang, Tian, Yuan, Zhang, Zicheng, Jia, Qi, Li, Chunyi, Zhang, Renrui, Li, Heng, Wang, Zongrui, Sun, Wei
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2601.19507
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917225207693312
author Zhu, Xiangyang
Tian, Yuan
Zhang, Zicheng
Jia, Qi
Li, Chunyi
Zhang, Renrui
Li, Heng
Wang, Zongrui
Sun, Wei
author_facet Zhu, Xiangyang
Tian, Yuan
Zhang, Zicheng
Jia, Qi
Li, Chunyi
Zhang, Renrui
Li, Heng
Wang, Zongrui
Sun, Wei
contents Large vision-language models (LVLMs) exhibit remarkable capabilities in cross-modal tasks but face significant safety challenges, which undermine their reliability in real-world applications. Efforts have been made to build LVLM safety evaluation benchmarks to uncover their vulnerability. However, existing benchmarks are hindered by their labor-intensive construction process, static complexity, and limited discriminative power. Thus, they may fail to keep pace with rapidly evolving models and emerging risks. To address these limitations, we propose VLSafetyBencher, the first automated system for LVLM safety benchmarking. VLSafetyBencher introduces four collaborative agents: Data Preprocessing, Generation, Augmentation, and Selection agents to construct and select high-quality samples. Experiments validates that VLSafetyBencher can construct high-quality safety benchmarks within one week at a minimal cost. The generated benchmark effectively distinguish safety, with a safety rate disparity of 70% between the most and least safe models.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19507
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs
Zhu, Xiangyang
Tian, Yuan
Zhang, Zicheng
Jia, Qi
Li, Chunyi
Zhang, Renrui
Li, Heng
Wang, Zongrui
Sun, Wei
Computation and Language
Large vision-language models (LVLMs) exhibit remarkable capabilities in cross-modal tasks but face significant safety challenges, which undermine their reliability in real-world applications. Efforts have been made to build LVLM safety evaluation benchmarks to uncover their vulnerability. However, existing benchmarks are hindered by their labor-intensive construction process, static complexity, and limited discriminative power. Thus, they may fail to keep pace with rapidly evolving models and emerging risks. To address these limitations, we propose VLSafetyBencher, the first automated system for LVLM safety benchmarking. VLSafetyBencher introduces four collaborative agents: Data Preprocessing, Generation, Augmentation, and Selection agents to construct and select high-quality samples. Experiments validates that VLSafetyBencher can construct high-quality safety benchmarks within one week at a minimal cost. The generated benchmark effectively distinguish safety, with a safety rate disparity of 70% between the most and least safe models.
title Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs
topic Computation and Language
url https://arxiv.org/abs/2601.19507