Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Xuannan, Li, Zekun, He, Zheqi, Li, Peipei, Xia, Shuhan, Cui, Xing, Huang, Huaibo, Yang, Xi, He, Ran
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918174221402112
author Liu, Xuannan
Li, Zekun
He, Zheqi
Li, Peipei
Xia, Shuhan
Cui, Xing
Huang, Huaibo
Yang, Xi
He, Ran
author_facet Liu, Xuannan
Li, Zekun
He, Zheqi
Li, Peipei
Xia, Shuhan
Cui, Xing
Huang, Huaibo
Yang, Xi
He, Ran
contents The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce distinct safety risks. To bridge this gap, we introduce Video-SafetyBench, the first comprehensive benchmark designed to evaluate the safety of LVLMs under video-text attacks. It comprises 2,264 video-text pairs spanning 48 fine-grained unsafe categories, each pairing a synthesized video with either a harmful query, which contains explicit malice, or a benign query, which appears harmless but triggers harmful behavior when interpreted alongside the video. To generate semantically accurate videos for safety evaluation, we design a controllable pipeline that decomposes video semantics into subject images (what is shown) and motion text (how it moves), which jointly guide the synthesis of query-relevant videos. To effectively evaluate uncertain or borderline harmful outputs, we propose RJScore, a novel LLM-based metric that incorporates the confidence of judge models and human-aligned decision threshold calibration. Extensive experiments show that benign-query video composition achieves average attack success rates of 67.2%, revealing consistent vulnerabilities to video-induced attacks. We believe Video-SafetyBench will catalyze future research into video-based safety evaluation and defense strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
Liu, Xuannan
Li, Zekun
He, Zheqi
Li, Peipei
Xia, Shuhan
Cui, Xing
Huang, Huaibo
Yang, Xi
He, Ran
Computer Vision and Pattern Recognition
Computation and Language
The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce distinct safety risks. To bridge this gap, we introduce Video-SafetyBench, the first comprehensive benchmark designed to evaluate the safety of LVLMs under video-text attacks. It comprises 2,264 video-text pairs spanning 48 fine-grained unsafe categories, each pairing a synthesized video with either a harmful query, which contains explicit malice, or a benign query, which appears harmless but triggers harmful behavior when interpreted alongside the video. To generate semantically accurate videos for safety evaluation, we design a controllable pipeline that decomposes video semantics into subject images (what is shown) and motion text (how it moves), which jointly guide the synthesis of query-relevant videos. To effectively evaluate uncertain or borderline harmful outputs, we propose RJScore, a novel LLM-based metric that incorporates the confidence of judge models and human-aligned decision threshold calibration. Extensive experiments show that benign-query video composition achieves average attack success rates of 67.2%, revealing consistent vulnerabilities to video-induced attacks. We believe Video-SafetyBench will catalyze future research into video-based safety evaluation and defense strategies.
title Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.11842