Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Xuannan, Cui, Xing, Li, Peipei, Li, Zekun, Huang, Huaibo, Xia, Shuhan, Zhang, Miaoxuan, Zou, Yueying, He, Ran
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915054535835648
author Liu, Xuannan
Cui, Xing
Li, Peipei
Li, Zekun
Huang, Huaibo
Xia, Shuhan
Zhang, Miaoxuan
Zou, Yueying
He, Ran
author_facet Liu, Xuannan
Cui, Xing
Li, Peipei
Li, Zekun
Huang, Huaibo
Xia, Shuhan
Zhang, Miaoxuan
Zou, Yueying
He, Ran
contents The rapid evolution of multimodal foundation models has led to significant advancements in cross-modal understanding and generation across diverse modalities, including text, images, audio, and video. However, these models remain susceptible to jailbreak attacks, which can bypass built-in safety mechanisms and induce the production of potentially harmful content. Consequently, understanding the methods of jailbreak attacks and existing defense mechanisms is essential to ensure the safe deployment of multimodal generative models in real-world scenarios, particularly in security-sensitive applications. To provide comprehensive insight into this topic, this survey reviews jailbreak and defense in multimodal generative models. First, given the generalized lifecycle of multimodal jailbreak, we systematically explore attacks and corresponding defense strategies across four levels: input, encoder, generator, and output. Based on this analysis, we present a detailed taxonomy of attack methods, defense mechanisms, and evaluation frameworks specific to multimodal generative models. Additionally, we cover a wide range of input-output configurations, including modalities such as Any-to-Text, Any-to-Vision, and Any-to-Any within generative systems. Finally, we highlight current research challenges and propose potential directions for future research. The open-source repository corresponding to this work can be found at https://github.com/liuxuannan/Awesome-Multimodal-Jailbreak.
format Preprint
id arxiv_https___arxiv_org_abs_2411_09259
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
Liu, Xuannan
Cui, Xing
Li, Peipei
Li, Zekun
Huang, Huaibo
Xia, Shuhan
Zhang, Miaoxuan
Zou, Yueying
He, Ran
Computer Vision and Pattern Recognition
Computation and Language
The rapid evolution of multimodal foundation models has led to significant advancements in cross-modal understanding and generation across diverse modalities, including text, images, audio, and video. However, these models remain susceptible to jailbreak attacks, which can bypass built-in safety mechanisms and induce the production of potentially harmful content. Consequently, understanding the methods of jailbreak attacks and existing defense mechanisms is essential to ensure the safe deployment of multimodal generative models in real-world scenarios, particularly in security-sensitive applications. To provide comprehensive insight into this topic, this survey reviews jailbreak and defense in multimodal generative models. First, given the generalized lifecycle of multimodal jailbreak, we systematically explore attacks and corresponding defense strategies across four levels: input, encoder, generator, and output. Based on this analysis, we present a detailed taxonomy of attack methods, defense mechanisms, and evaluation frameworks specific to multimodal generative models. Additionally, we cover a wide range of input-output configurations, including modalities such as Any-to-Text, Any-to-Vision, and Any-to-Any within generative systems. Finally, we highlight current research challenges and propose potential directions for future research. The open-source repository corresponding to this work can be found at https://github.com/liuxuannan/Awesome-Multimodal-Jailbreak.
title Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2411.09259