UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Segyu, Cho, Boryeong, Jung, Hojung, An, Seokhyun, Kim, Juhyeong, Kwak, Jaehyun, Yang, Yongjin, Jang, Sangwon, Park, Youngrok, Chang, Wonjun, Yun, Se-Young
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910057849946112
author Lee, Segyu
Cho, Boryeong
Jung, Hojung
An, Seokhyun
Kim, Juhyeong
Kwak, Jaehyun
Yang, Yongjin
Jang, Sangwon
Park, Youngrok
Chang, Wonjun
Yun, Se-Young
author_facet Lee, Segyu
Cho, Boryeong
Jung, Hojung
An, Seokhyun
Kim, Juhyeong
Kwak, Jaehyun
Yang, Yongjin
Jang, Sangwon
Park, Youngrok
Chang, Wonjun
Yun, Se-Young
contents Unified Multimodal Models (UMMs) offer powerful cross-modality capabilities but introduce new safety risks not observed in single-task models. Despite their emergence, existing safety benchmarks remain fragmented across tasks and modalities, limiting the comprehensive evaluation of complex system-level vulnerabilities. To address this gap, we introduce UniSAFE, the first comprehensive benchmark for system-level safety evaluation of UMMs across 7 I/O modality combinations, spanning conventional tasks and novel multimodal-context image generation settings. UniSAFE is built with a shared-target design that projects common risk scenarios across task-specific I/O configurations, enabling controlled cross-task comparisons of safety failures. Comprising 6,802 curated instances, we use UniSAFE to evaluate 15 state-of-the-art UMMs, both proprietary and open-source. Our results reveal critical vulnerabilities across current UMMs, including elevated safety violations in multi-image composition and multi-turn settings, with image-output tasks consistently more vulnerable than text-output tasks. These findings highlight the need for stronger system-level safety alignment for UMMs. Our code and data are publicly available at https://github.com/segyulee/UniSAFE
format Preprint
id arxiv_https___arxiv_org_abs_2603_17476
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
Lee, Segyu
Cho, Boryeong
Jung, Hojung
An, Seokhyun
Kim, Juhyeong
Kwak, Jaehyun
Yang, Yongjin
Jang, Sangwon
Park, Youngrok
Chang, Wonjun
Yun, Se-Young
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Unified Multimodal Models (UMMs) offer powerful cross-modality capabilities but introduce new safety risks not observed in single-task models. Despite their emergence, existing safety benchmarks remain fragmented across tasks and modalities, limiting the comprehensive evaluation of complex system-level vulnerabilities. To address this gap, we introduce UniSAFE, the first comprehensive benchmark for system-level safety evaluation of UMMs across 7 I/O modality combinations, spanning conventional tasks and novel multimodal-context image generation settings. UniSAFE is built with a shared-target design that projects common risk scenarios across task-specific I/O configurations, enabling controlled cross-task comparisons of safety failures. Comprising 6,802 curated instances, we use UniSAFE to evaluate 15 state-of-the-art UMMs, both proprietary and open-source. Our results reveal critical vulnerabilities across current UMMs, including elevated safety violations in multi-image composition and multi-turn settings, with image-output tasks consistently more vulnerable than text-output tasks. These findings highlight the need for stronger system-level safety alignment for UMMs. Our code and data are publicly available at https://github.com/segyulee/UniSAFE
title UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.17476