Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Peng-Fei, Huang, Zi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911451709440000
author Zhang, Peng-Fei
Huang, Zi
author_facet Zhang, Peng-Fei
Huang, Zi
contents Existing adversarial attacks for VLP models are mostly sample-specific, resulting in substantial computational overhead when scaled to large datasets or new scenarios. To overcome this limitation, we propose Hierarchical Refinement Attack (HRA), a multimodal universal attack framework for VLP models. For the image modality, we refine the optimization path by leveraging a temporal hierarchy of historical and estimated future gradients to avoid local minima and stabilize universal perturbation learning. For the text modality, it hierarchically models textual importance by considering both intra- and inter-sentence contributions to identify globally influential words, which are then used as universal text perturbations. Extensive experiments across various downstream tasks, VLP models, and datasets, demonstrate the superior transferability of the proposed universal multimodal attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10313
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
Zhang, Peng-Fei
Huang, Zi
Computer Vision and Pattern Recognition
Multimedia
Existing adversarial attacks for VLP models are mostly sample-specific, resulting in substantial computational overhead when scaled to large datasets or new scenarios. To overcome this limitation, we propose Hierarchical Refinement Attack (HRA), a multimodal universal attack framework for VLP models. For the image modality, we refine the optimization path by leveraging a temporal hierarchy of historical and estimated future gradients to avoid local minima and stabilize universal perturbation learning. For the text modality, it hierarchically models textual importance by considering both intra- and inter-sentence contributions to identify globally influential words, which are then used as universal text perturbations. Extensive experiments across various downstream tasks, VLP models, and datasets, demonstrate the superior transferability of the proposed universal multimodal attacks.
title Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2601.10313