Where and How to Attack? A Causality-Inspired Recipe for Generating Counterfactual Adversarial Examples

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Ruichu, Zhu, Yuxuan, Qiao, Jie, Liang, Zefeng, Liu, Furui, Hao, Zhifeng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913210792148992
author Cai, Ruichu
Zhu, Yuxuan
Qiao, Jie
Liang, Zefeng
Liu, Furui
Hao, Zhifeng
author_facet Cai, Ruichu
Zhu, Yuxuan
Qiao, Jie
Liang, Zefeng
Liu, Furui
Hao, Zhifeng
contents Deep neural networks (DNNs) have been demonstrated to be vulnerable to well-crafted \emph{adversarial examples}, which are generated through either well-conceived $\mathcal{L}_p$-norm restricted or unrestricted attacks. Nevertheless, the majority of those approaches assume that adversaries can modify any features as they wish, and neglect the causal generating process of the data, which is unreasonable and unpractical. For instance, a modification in income would inevitably impact features like the debt-to-income ratio within a banking system. By considering the underappreciated causal generating process, first, we pinpoint the source of the vulnerability of DNNs via the lens of causality, then give theoretical results to answer \emph{where to attack}. Second, considering the consequences of the attack interventions on the current state of the examples to generate more realistic adversarial examples, we propose CADE, a framework that can generate \textbf{C}ounterfactual \textbf{AD}versarial \textbf{E}xamples to answer \emph{how to attack}. The empirical results demonstrate CADE's effectiveness, as evidenced by its competitive performance across diverse attack scenarios, including white-box, transfer-based, and random intervention attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2312_13628
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Where and How to Attack? A Causality-Inspired Recipe for Generating Counterfactual Adversarial Examples
Cai, Ruichu
Zhu, Yuxuan
Qiao, Jie
Liang, Zefeng
Liu, Furui
Hao, Zhifeng
Machine Learning
Deep neural networks (DNNs) have been demonstrated to be vulnerable to well-crafted \emph{adversarial examples}, which are generated through either well-conceived $\mathcal{L}_p$-norm restricted or unrestricted attacks. Nevertheless, the majority of those approaches assume that adversaries can modify any features as they wish, and neglect the causal generating process of the data, which is unreasonable and unpractical. For instance, a modification in income would inevitably impact features like the debt-to-income ratio within a banking system. By considering the underappreciated causal generating process, first, we pinpoint the source of the vulnerability of DNNs via the lens of causality, then give theoretical results to answer \emph{where to attack}. Second, considering the consequences of the attack interventions on the current state of the examples to generate more realistic adversarial examples, we propose CADE, a framework that can generate \textbf{C}ounterfactual \textbf{AD}versarial \textbf{E}xamples to answer \emph{how to attack}. The empirical results demonstrate CADE's effectiveness, as evidenced by its competitive performance across diverse attack scenarios, including white-box, transfer-based, and random intervention attacks.
title Where and How to Attack? A Causality-Inspired Recipe for Generating Counterfactual Adversarial Examples
topic Machine Learning
url https://arxiv.org/abs/2312.13628