MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Chejian, Zhang, Jiawei, Chen, Zhaorun, Xie, Chulin, Kang, Mintong, Potter, Yujin, Wang, Zhun, Yuan, Zhuowen, Xiong, Alexander, Xiong, Zidi, Zhang, Chenhui, Yuan, Lingzhi, Zeng, Yi, Xu, Peiyang, Guo, Chengquan, Zhou, Andy, Tan, Jeffrey Ziwei, Zhao, Xuandong, Pinto, Francesco, Xiang, Zhen, Gai, Yu, Lin, Zinan, Hendrycks, Dan, Li, Bo, Song, Dawn
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915204385734656
author Xu, Chejian
Zhang, Jiawei
Chen, Zhaorun
Xie, Chulin
Kang, Mintong
Potter, Yujin
Wang, Zhun
Yuan, Zhuowen
Xiong, Alexander
Xiong, Zidi
Zhang, Chenhui
Yuan, Lingzhi
Zeng, Yi
Xu, Peiyang
Guo, Chengquan
Zhou, Andy
Tan, Jeffrey Ziwei
Zhao, Xuandong
Pinto, Francesco
Xiang, Zhen
Gai, Yu
Lin, Zinan
Hendrycks, Dan
Li, Bo
Song, Dawn
author_facet Xu, Chejian
Zhang, Jiawei
Chen, Zhaorun
Xie, Chulin
Kang, Mintong
Potter, Yujin
Wang, Zhun
Yuan, Zhuowen
Xiong, Alexander
Xiong, Zidi
Zhang, Chenhui
Yuan, Lingzhi
Zeng, Yi
Xu, Peiyang
Guo, Chengquan
Zhou, Andy
Tan, Jeffrey Ziwei
Zhao, Xuandong
Pinto, Francesco
Xiang, Zhen
Gai, Yu
Lin, Zinan
Hendrycks, Dan
Li, Bo
Song, Dawn
contents Multimodal foundation models (MMFMs) play a crucial role in various applications, including autonomous driving, healthcare, and virtual assistants. However, several studies have revealed vulnerabilities in these models, such as generating unsafe content by text-to-image models. Existing benchmarks on multimodal models either predominantly assess the helpfulness of these models, or only focus on limited perspectives such as fairness and privacy. In this paper, we present the first unified platform, MMDT (Multimodal DecodingTrust), designed to provide a comprehensive safety and trustworthiness evaluation for MMFMs. Our platform assesses models from multiple perspectives, including safety, hallucination, fairness/bias, privacy, adversarial robustness, and out-of-distribution (OOD) generalization. We have designed various evaluation scenarios and red teaming algorithms under different tasks for each perspective to generate challenging data, forming a high-quality benchmark. We evaluate a range of multimodal models using MMDT, and our findings reveal a series of vulnerabilities and areas for improvement across these perspectives. This work introduces the first comprehensive and unique safety and trustworthiness evaluation platform for MMFMs, paving the way for developing safer and more reliable MMFMs and systems. Our platform and benchmark are available at https://mmdecodingtrust.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
Xu, Chejian
Zhang, Jiawei
Chen, Zhaorun
Xie, Chulin
Kang, Mintong
Potter, Yujin
Wang, Zhun
Yuan, Zhuowen
Xiong, Alexander
Xiong, Zidi
Zhang, Chenhui
Yuan, Lingzhi
Zeng, Yi
Xu, Peiyang
Guo, Chengquan
Zhou, Andy
Tan, Jeffrey Ziwei
Zhao, Xuandong
Pinto, Francesco
Xiang, Zhen
Gai, Yu
Lin, Zinan
Hendrycks, Dan
Li, Bo
Song, Dawn
Computation and Language
Artificial Intelligence
Cryptography and Security
Multimodal foundation models (MMFMs) play a crucial role in various applications, including autonomous driving, healthcare, and virtual assistants. However, several studies have revealed vulnerabilities in these models, such as generating unsafe content by text-to-image models. Existing benchmarks on multimodal models either predominantly assess the helpfulness of these models, or only focus on limited perspectives such as fairness and privacy. In this paper, we present the first unified platform, MMDT (Multimodal DecodingTrust), designed to provide a comprehensive safety and trustworthiness evaluation for MMFMs. Our platform assesses models from multiple perspectives, including safety, hallucination, fairness/bias, privacy, adversarial robustness, and out-of-distribution (OOD) generalization. We have designed various evaluation scenarios and red teaming algorithms under different tasks for each perspective to generate challenging data, forming a high-quality benchmark. We evaluate a range of multimodal models using MMDT, and our findings reveal a series of vulnerabilities and areas for improvement across these perspectives. This work introduces the first comprehensive and unique safety and trustworthiness evaluation platform for MMFMs, paving the way for developing safer and more reliable MMFMs and systems. Our platform and benchmark are available at https://mmdecodingtrust.github.io/.
title MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
topic Computation and Language
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2503.14827