Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Jiani, Wang, Zhiyuan, Zhang, Zeliang, Huang, Chao, Liang, Susan, Tang, Yunlong, Xu, Chenliang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912711514783744
author Liu, Jiani
Wang, Zhiyuan
Zhang, Zeliang
Huang, Chao
Liang, Susan
Tang, Yunlong
Xu, Chenliang
author_facet Liu, Jiani
Wang, Zhiyuan
Zhang, Zeliang
Huang, Chao
Liang, Susan
Tang, Yunlong
Xu, Chenliang
contents Vision Transformers (ViTs) have demonstrated impressive performance across a range of applications, including many safety-critical tasks. However, their unique architectural properties raise new challenges and opportunities in adversarial robustness. In particular, we observe that adversarial examples crafted on ViTs exhibit higher transferability compared to those crafted on CNNs, suggesting that ViTs contain structural characteristics favorable for transferable attacks. In this work, we investigate the role of computational redundancy in ViTs and its impact on adversarial transferability. Unlike prior studies that aim to reduce computation for efficiency, we propose to exploit this redundancy to improve the quality and transferability of adversarial examples. Through a detailed analysis, we identify two forms of redundancy, including the data-level and model-level, that can be harnessed to amplify attack effectiveness. Building on this insight, we design a suite of techniques, including attention sparsity manipulation, attention head permutation, clean token regularization, ghost MoE diversification, and test-time adversarial training. Extensive experiments on the ImageNet-1k dataset validate the effectiveness of our approach, showing that our methods significantly outperform existing baselines in both transferability and generality across diverse model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10804
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
Liu, Jiani
Wang, Zhiyuan
Zhang, Zeliang
Huang, Chao
Liang, Susan
Tang, Yunlong
Xu, Chenliang
Computer Vision and Pattern Recognition
Vision Transformers (ViTs) have demonstrated impressive performance across a range of applications, including many safety-critical tasks. However, their unique architectural properties raise new challenges and opportunities in adversarial robustness. In particular, we observe that adversarial examples crafted on ViTs exhibit higher transferability compared to those crafted on CNNs, suggesting that ViTs contain structural characteristics favorable for transferable attacks. In this work, we investigate the role of computational redundancy in ViTs and its impact on adversarial transferability. Unlike prior studies that aim to reduce computation for efficiency, we propose to exploit this redundancy to improve the quality and transferability of adversarial examples. Through a detailed analysis, we identify two forms of redundancy, including the data-level and model-level, that can be harnessed to amplify attack effectiveness. Building on this insight, we design a suite of techniques, including attention sparsity manipulation, attention head permutation, clean token regularization, ghost MoE diversification, and test-time adversarial training. Extensive experiments on the ImageNet-1k dataset validate the effectiveness of our approach, showing that our methods significantly outperform existing baselines in both transferability and generality across diverse model architectures.
title Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.10804