Saved in:
Bibliographic Details
Main Authors: Peng, Zixiang, Xu, Yongxiu, Zhang, Qinyi, Shen, Jiexun, Zhang, Yifan, Xu, Hongbo, Wang, Yubin, Gou, Gaopeng
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.00547
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918423046389760
author Peng, Zixiang
Xu, Yongxiu
Zhang, Qinyi
Shen, Jiexun
Zhang, Yifan
Xu, Hongbo
Wang, Yubin
Gou, Gaopeng
author_facet Peng, Zixiang
Xu, Yongxiu
Zhang, Qinyi
Shen, Jiexun
Zhang, Yifan
Xu, Hongbo
Wang, Yubin
Gou, Gaopeng
contents Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. While this architectural unification, driven by the deep fusion of multimodal features, enhances model performance, it also introduces important yet underexplored safety challenges. Existing safety benchmarks predominantly focus on isolated understanding or generation tasks, failing to evaluate the holistic safety of UMLMs when handling diverse tasks under a unified framework. To address this, we introduce Uni-SafeBench, a comprehensive benchmark featuring a taxonomy of six major safety categories across seven task types. To ensure rigorous assessment, we develop Uni-Judger, a framework that effectively decouples contextual safety from intrinsic safety. Based on comprehensive evaluations across Uni-SafeBench, we uncover that while the unification process enhances model capabilities, it significantly degrades the inherent safety of the underlying LLM. Furthermore, open-source UMLMs exhibit much lower safety performance than multimodal large models specialized for either generation or understanding tasks. We open-source all resources to systematically expose these risks and foster safer AGI development.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00547
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models
Peng, Zixiang
Xu, Yongxiu
Zhang, Qinyi
Shen, Jiexun
Zhang, Yifan
Xu, Hongbo
Wang, Yubin
Gou, Gaopeng
Artificial Intelligence
Machine Learning
Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. While this architectural unification, driven by the deep fusion of multimodal features, enhances model performance, it also introduces important yet underexplored safety challenges. Existing safety benchmarks predominantly focus on isolated understanding or generation tasks, failing to evaluate the holistic safety of UMLMs when handling diverse tasks under a unified framework. To address this, we introduce Uni-SafeBench, a comprehensive benchmark featuring a taxonomy of six major safety categories across seven task types. To ensure rigorous assessment, we develop Uni-Judger, a framework that effectively decouples contextual safety from intrinsic safety. Based on comprehensive evaluations across Uni-SafeBench, we uncover that while the unification process enhances model capabilities, it significantly degrades the inherent safety of the underlying LLM. Furthermore, open-source UMLMs exhibit much lower safety performance than multimodal large models specialized for either generation or understanding tasks. We open-source all resources to systematically expose these risks and foster safer AGI development.
title Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.00547