Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zhiqiang, Liu, Dongrui, Li, Yan, Ying, Zonghao, Xue, Wei, Luo, Wenhan, Guo, Yike
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917505420754944
author Wang, Zhiqiang
Liu, Dongrui
Li, Yan
Ying, Zonghao
Xue, Wei
Luo, Wenhan
Guo, Yike
author_facet Wang, Zhiqiang
Liu, Dongrui
Li, Yan
Ying, Zonghao
Xue, Wei
Luo, Wenhan
Guo, Yike
contents Existing adversarial attacks on vision-language models (VLMs) can steer model outputs toward attacker-specified target responses, but their effectiveness often degrades when the same perturbed input is paired with different textual queries. This paper studies cross-query response manipulation, where a single adversarial example is expected to remain effective across diverse user queries. We first analyze the limitations of existing attacks and find that successful transfer is closely associated with preserving an image-dominant attention pattern during response generation. Motivated by the observation, we propose \textbf{Attention Hijacking}, a novel adversarial attack that explicitly steers internal attention distributions toward a persistent image-dominant pattern. By amplifying the influence of visual tokens on target response tokens while suppressing the competing influence of textual tokens, our method reduces the dependence of the manipulated output on the specific wording of the query. Extensive experiments on widely used VLMs show that Attention Hijacking substantially improves cross-query transferability across diverse target responses and unseen queries. The method also extends effectively to multiple attack scenarios, offering new insights into the role of attention stability in transferable response manipulation for VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17310
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
Wang, Zhiqiang
Liu, Dongrui
Li, Yan
Ying, Zonghao
Xue, Wei
Luo, Wenhan
Guo, Yike
Computer Vision and Pattern Recognition
Artificial Intelligence
Existing adversarial attacks on vision-language models (VLMs) can steer model outputs toward attacker-specified target responses, but their effectiveness often degrades when the same perturbed input is paired with different textual queries. This paper studies cross-query response manipulation, where a single adversarial example is expected to remain effective across diverse user queries. We first analyze the limitations of existing attacks and find that successful transfer is closely associated with preserving an image-dominant attention pattern during response generation. Motivated by the observation, we propose \textbf{Attention Hijacking}, a novel adversarial attack that explicitly steers internal attention distributions toward a persistent image-dominant pattern. By amplifying the influence of visual tokens on target response tokens while suppressing the competing influence of textual tokens, our method reduces the dependence of the manipulated output on the specific wording of the query. Extensive experiments on widely used VLMs show that Attention Hijacking substantially improves cross-query transferability across diverse target responses and unseen queries. The method also extends effectively to multiple attack scenarios, offering new insights into the role of attention stability in transferable response manipulation for VLMs.
title Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.17310