Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Hao, Xiao, Erjia, Yang, Jiayan, Duan, Jinhao, Wang, Yichi, Cao, Jiahang, Zhang, Qiang, Yang, Le, Xu, Kaidi, Gu, Jindong, Xu, Renjing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913949030547456
author Cheng, Hao
Xiao, Erjia
Yang, Jiayan
Duan, Jinhao
Wang, Yichi
Cao, Jiahang
Zhang, Qiang
Yang, Le
Xu, Kaidi
Gu, Jindong
Xu, Renjing
author_facet Cheng, Hao
Xiao, Erjia
Yang, Jiayan
Duan, Jinhao
Wang, Yichi
Cao, Jiahang
Zhang, Qiang
Yang, Le
Xu, Kaidi
Gu, Jindong
Xu, Renjing
contents Multimodal Large Language Models (MLLMs) demonstrate exceptional performance in cross-modality interaction, yet they also suffer adversarial vulnerabilities. In particular, the transferability of adversarial examples remains an ongoing challenge. In this paper, we specifically analyze the manifestation of adversarial transferability among MLLMs and identify the key factors that influence this characteristic. We discover that the transferability of MLLMs exists in cross-LLM scenarios with the same vision encoder and indicate \underline{\textit{two key Factors}} that may influence transferability. We provide two semantic-level data augmentation methods, Adding Image Patch (AIP) and Typography Augment Transferability Method (TATM), which boost the transferability of adversarial examples across MLLMs. To explore the potential impact in the real world, we utilize two tasks that can have both negative and positive societal impacts: \ding{182} Harmful Content Insertion and \ding{183} Information Protection.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20090
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models
Cheng, Hao
Xiao, Erjia
Yang, Jiayan
Duan, Jinhao
Wang, Yichi
Cao, Jiahang
Zhang, Qiang
Yang, Le
Xu, Kaidi
Gu, Jindong
Xu, Renjing
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) demonstrate exceptional performance in cross-modality interaction, yet they also suffer adversarial vulnerabilities. In particular, the transferability of adversarial examples remains an ongoing challenge. In this paper, we specifically analyze the manifestation of adversarial transferability among MLLMs and identify the key factors that influence this characteristic. We discover that the transferability of MLLMs exists in cross-LLM scenarios with the same vision encoder and indicate \underline{\textit{two key Factors}} that may influence transferability. We provide two semantic-level data augmentation methods, Adding Image Patch (AIP) and Typography Augment Transferability Method (TATM), which boost the transferability of adversarial examples across MLLMs. To explore the potential impact in the real world, we utilize two tasks that can have both negative and positive societal impacts: \ding{182} Harmful Content Insertion and \ding{183} Information Protection.
title Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.20090