Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ren, Yuchen, Zhao, Zhengyu, Lin, Chenhao, Yang, Bo, Zhou, Lu, Liu, Zhe, Shen, Chao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915080764915712
author Ren, Yuchen
Zhao, Zhengyu
Lin, Chenhao
Yang, Bo
Zhou, Lu
Liu, Zhe
Shen, Chao
author_facet Ren, Yuchen
Zhao, Zhengyu
Lin, Chenhao
Yang, Bo
Zhou, Lu
Liu, Zhe
Shen, Chao
contents Transferable adversarial examples are known to cause threats in practical, black-box attack scenarios. A notable approach to improving transferability is using integrated gradients (IG), originally developed for model interpretability. In this paper, we find that existing IG-based attacks have limited transferability due to their naive adoption of IG in model interpretability. To address this limitation, we focus on the IG integration path and refine it in three aspects: multiplicity, monotonicity, and diversity, supported by theoretical analyses. We propose the Multiple Monotonic Diversified Integrated Gradients (MuMoDIG) attack, which can generate highly transferable adversarial examples on different CNN and ViT models and defenses. Experiments validate that MuMoDIG outperforms the latest IG-based attack by up to 37.3\% and other state-of-the-art attacks by 8.4\%. In general, our study reveals that migrating established techniques to improve transferability may require non-trivial efforts. Code is available at \url{https://github.com/RYC-98/MuMoDIG}.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18844
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path
Ren, Yuchen
Zhao, Zhengyu
Lin, Chenhao
Yang, Bo
Zhou, Lu
Liu, Zhe
Shen, Chao
Cryptography and Security
Computer Vision and Pattern Recognition
Machine Learning
Transferable adversarial examples are known to cause threats in practical, black-box attack scenarios. A notable approach to improving transferability is using integrated gradients (IG), originally developed for model interpretability. In this paper, we find that existing IG-based attacks have limited transferability due to their naive adoption of IG in model interpretability. To address this limitation, we focus on the IG integration path and refine it in three aspects: multiplicity, monotonicity, and diversity, supported by theoretical analyses. We propose the Multiple Monotonic Diversified Integrated Gradients (MuMoDIG) attack, which can generate highly transferable adversarial examples on different CNN and ViT models and defenses. Experiments validate that MuMoDIG outperforms the latest IG-based attack by up to 37.3\% and other state-of-the-art attacks by 8.4\%. In general, our study reveals that migrating established techniques to improve transferability may require non-trivial efforts. Code is available at \url{https://github.com/RYC-98/MuMoDIG}.
title Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path
topic Cryptography and Security
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.18844