SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hanqing, Tian, Yuan, Liu, Mingyu, Zhang, Zhenhao, Zhu, Xiangyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915705847283712
author Wang, Hanqing
Tian, Yuan
Liu, Mingyu
Zhang, Zhenhao
Zhu, Xiangyang
author_facet Wang, Hanqing
Tian, Yuan
Liu, Mingyu
Zhang, Zhenhao
Zhu, Xiangyang
contents In the rapidly evolving landscape of Multimodal Large Language Models (MLLMs), the safety concerns of their outputs have earned significant attention. Although numerous datasets have been proposed, they may become outdated with MLLM advancements and are susceptible to data contamination issues. To address these problems, we propose \textbf{SDEval}, the \textit{first} safety dynamic evaluation framework to controllably adjust the distribution and complexity of safety benchmarks. Specifically, SDEval mainly adopts three dynamic strategies: text, image, and text-image dynamics to generate new samples from original benchmarks. We first explore the individual effects of text and image dynamics on model safety. Then, we find that injecting text dynamics into images can further impact safety, and conversely, injecting image dynamics into text also leads to safety risks. SDEval is general enough to be applied to various existing safety and even capability benchmarks. Experiments across safety benchmarks, MLLMGuard and VLSBench, and capability benchmarks, MMBench and MMVet, show that SDEval significantly influences safety evaluation, mitigates data contamination, and exposes safety limitations of MLLMs. Code is available at https://github.com/hq-King/SDEval
format Preprint
id arxiv_https___arxiv_org_abs_2508_06142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
Wang, Hanqing
Tian, Yuan
Liu, Mingyu
Zhang, Zhenhao
Zhu, Xiangyang
Computer Vision and Pattern Recognition
In the rapidly evolving landscape of Multimodal Large Language Models (MLLMs), the safety concerns of their outputs have earned significant attention. Although numerous datasets have been proposed, they may become outdated with MLLM advancements and are susceptible to data contamination issues. To address these problems, we propose \textbf{SDEval}, the \textit{first} safety dynamic evaluation framework to controllably adjust the distribution and complexity of safety benchmarks. Specifically, SDEval mainly adopts three dynamic strategies: text, image, and text-image dynamics to generate new samples from original benchmarks. We first explore the individual effects of text and image dynamics on model safety. Then, we find that injecting text dynamics into images can further impact safety, and conversely, injecting image dynamics into text also leads to safety risks. SDEval is general enough to be applied to various existing safety and even capability benchmarks. Experiments across safety benchmarks, MLLMGuard and VLSBench, and capability benchmarks, MMBench and MMVet, show that SDEval significantly influences safety evaluation, mitigates data contamination, and exposes safety limitations of MLLMs. Code is available at https://github.com/hq-King/SDEval
title SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.06142