Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Wenxuan, Liu, Xiaoyuan, Gao, Kuiyi, Huang, Jen-tse, Yuan, Youliang, He, Pinjia, Wang, Shuai, Tu, Zhaopeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910983760379904
author Wang, Wenxuan
Liu, Xiaoyuan
Gao, Kuiyi
Huang, Jen-tse
Yuan, Youliang
He, Pinjia
Wang, Shuai
Tu, Zhaopeng
author_facet Wang, Wenxuan
Liu, Xiaoyuan
Gao, Kuiyi
Huang, Jen-tse
Yuan, Youliang
He, Pinjia
Wang, Shuai
Tu, Zhaopeng
contents Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the safety of these models remains a significant challenge, particularly in accurately identifying whether multimodal content is safe or unsafe-a capability we term safety awareness. In this paper, we introduce MMSafeAware, the first comprehensive multimodal safety awareness benchmark designed to evaluate MLLMs across 29 safety scenarios with 1500 carefully curated image-prompt pairs. MMSafeAware includes both unsafe and over-safety subsets to assess models abilities to correctly identify unsafe content and avoid over-sensitivity that can hinder helpfulness. Evaluating nine widely used MLLMs using MMSafeAware reveals that current models are not sufficiently safe and often overly sensitive; for example, GPT-4V misclassifies 36.1% of unsafe inputs as safe and 59.9% of benign inputs as unsafe. We further explore three methods to improve safety awareness-prompting-based approaches, visual contrastive decoding, and vision-centric reasoning fine-tuning-but find that none achieve satisfactory performance. Our findings highlight the profound challenges in developing MLLMs with robust safety awareness, underscoring the need for further research in this area. All the code and data will be publicly available to facilitate future research.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
Wang, Wenxuan
Liu, Xiaoyuan
Gao, Kuiyi
Huang, Jen-tse
Yuan, Youliang
He, Pinjia
Wang, Shuai
Tu, Zhaopeng
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the safety of these models remains a significant challenge, particularly in accurately identifying whether multimodal content is safe or unsafe-a capability we term safety awareness. In this paper, we introduce MMSafeAware, the first comprehensive multimodal safety awareness benchmark designed to evaluate MLLMs across 29 safety scenarios with 1500 carefully curated image-prompt pairs. MMSafeAware includes both unsafe and over-safety subsets to assess models abilities to correctly identify unsafe content and avoid over-sensitivity that can hinder helpfulness. Evaluating nine widely used MLLMs using MMSafeAware reveals that current models are not sufficiently safe and often overly sensitive; for example, GPT-4V misclassifies 36.1% of unsafe inputs as safe and 59.9% of benign inputs as unsafe. We further explore three methods to improve safety awareness-prompting-based approaches, visual contrastive decoding, and vision-centric reasoning fine-tuning-but find that none achieve satisfactory performance. Our findings highlight the profound challenges in developing MLLMs with robust safety awareness, underscoring the need for further research in this area. All the code and data will be publicly available to facilitate future research.
title Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2502.11184