On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xinwei, Liu, Hangcheng, Bai, Li, Wang, Hao, Ye, Qingqing, Zhang, Tianwei, Hu, Haibo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917504140443648
author Zhang, Xinwei
Liu, Hangcheng
Bai, Li
Wang, Hao
Ye, Qingqing
Zhang, Tianwei
Hu, Haibo
author_facet Zhang, Xinwei
Liu, Hangcheng
Bai, Li
Wang, Hao
Ye, Qingqing
Zhang, Tianwei
Hu, Haibo
contents Visual token compression is widely used to accelerate large vision-language models (LVLMs) by pruning or merging visual tokens, yet its adversarial robustness remains unexplored. We show that existing encoder-based attacks cannot fully disclose the robustness vulnerabilities of compressed LVLMs, due to an optimization-inference mismatch: perturbations are optimized on the full-token representation, while inference is performed through a token-compression bottleneck. To address this gap, we propose the Compression-AliGnEd attack (CAGE), which aligns perturbation optimization with compression inference without assuming access to the deployed compression mechanism or its token budget. CAGE combines (i) expected feature disruption, which concentrates distortion on tokens likely to survive across plausible budgets, and (ii) rank distortion alignment, which actively aligns token distortions with rank scores to promote the retention of highly distorted evidence. Across diverse representative plug-and-play compression mechanisms and datasets, our results show that CAGE consistently achieves lower robust accuracy than the baseline. This work highlights that robustness assessments ignoring compression can be overly optimistic, calling for compression-aware security evaluation and defenses for efficient LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21531
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
Zhang, Xinwei
Liu, Hangcheng
Bai, Li
Wang, Hao
Ye, Qingqing
Zhang, Tianwei
Hu, Haibo
Cryptography and Security
Artificial Intelligence
Computer Vision and Pattern Recognition
Visual token compression is widely used to accelerate large vision-language models (LVLMs) by pruning or merging visual tokens, yet its adversarial robustness remains unexplored. We show that existing encoder-based attacks cannot fully disclose the robustness vulnerabilities of compressed LVLMs, due to an optimization-inference mismatch: perturbations are optimized on the full-token representation, while inference is performed through a token-compression bottleneck. To address this gap, we propose the Compression-AliGnEd attack (CAGE), which aligns perturbation optimization with compression inference without assuming access to the deployed compression mechanism or its token budget. CAGE combines (i) expected feature disruption, which concentrates distortion on tokens likely to survive across plausible budgets, and (ii) rank distortion alignment, which actively aligns token distortions with rank scores to promote the retention of highly distorted evidence. Across diverse representative plug-and-play compression mechanisms and datasets, our results show that CAGE consistently achieves lower robust accuracy than the baseline. This work highlights that robustness assessments ignoring compression can be overly optimistic, calling for compression-aware security evaluation and defenses for efficient LVLMs.
title On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
topic Cryptography and Security
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.21531