Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Chenhe, Gu, Jindong, Hua, Andong, Qin, Yao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915175532068864
author Gu, Chenhe
Gu, Jindong
Hua, Andong
Qin, Yao
author_facet Gu, Chenhe
Gu, Jindong
Hua, Andong
Qin, Yao
contents Multimodal Large Language Models (MLLMs), built upon LLMs, have recently gained attention for their capabilities in image recognition and understanding. However, while MLLMs are vulnerable to adversarial attacks, the transferability of these attacks across different models remains limited, especially under targeted attack setting. Existing methods primarily focus on vision-specific perturbations but struggle with the complex nature of vision-language modality alignment. In this work, we introduce the Dynamic Vision-Language Alignment (DynVLA) Attack, a novel approach that injects dynamic perturbations into the vision-language connector to enhance generalization across diverse vision-language alignment of different models. Our experimental results show that DynVLA significantly improves the transferability of adversarial examples across various MLLMs, including BLIP2, InstructBLIP, MiniGPT4, LLaVA, and closed-source models such as Gemini.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19672
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack
Gu, Chenhe
Gu, Jindong
Hua, Andong
Qin, Yao
Computer Vision and Pattern Recognition
Machine Learning
Multimodal Large Language Models (MLLMs), built upon LLMs, have recently gained attention for their capabilities in image recognition and understanding. However, while MLLMs are vulnerable to adversarial attacks, the transferability of these attacks across different models remains limited, especially under targeted attack setting. Existing methods primarily focus on vision-specific perturbations but struggle with the complex nature of vision-language modality alignment. In this work, we introduce the Dynamic Vision-Language Alignment (DynVLA) Attack, a novel approach that injects dynamic perturbations into the vision-language connector to enhance generalization across diverse vision-language alignment of different models. Our experimental results show that DynVLA significantly improves the transferability of adversarial examples across various MLLMs, including BLIP2, InstructBLIP, MiniGPT4, LLaVA, and closed-source models such as Gemini.
title Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2502.19672