Causal Interpretability for Adversarial Robustness: A Hybrid Generative Classification Approach

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Chunheng, Pisu, Pierluigi, Comert, Gurcan, Begashaw, Negash, Vaidyan, Varghese, Hubig, Nina Christine
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912752074752000
author Zhao, Chunheng
Pisu, Pierluigi
Comert, Gurcan
Begashaw, Negash
Vaidyan, Varghese
Hubig, Nina Christine
author_facet Zhao, Chunheng
Pisu, Pierluigi
Comert, Gurcan
Begashaw, Negash
Vaidyan, Varghese
Hubig, Nina Christine
contents Deep learning-based discriminative classifiers, despite their remarkable success, remain vulnerable to adversarial examples that can mislead model predictions. While adversarial training can enhance robustness, it fails to address the intrinsic vulnerability stemming from the opaque nature of these black-box models. We present a deep ensemble model that combines discriminative features with generative models to achieve both high accuracy and adversarial robustness. Our approach integrates a bottom-level pre-trained discriminative network for feature extraction with a top-level generative classification network that models adversarial input distributions through a deep latent variable model. Using variational Bayes, our model achieves superior robustness against white-box adversarial attacks without adversarial training. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate our model's superior adversarial robustness. Through evaluations using counterfactual metrics and feature interaction-based metrics, we establish correlations between model interpretability and adversarial robustness. Additionally, preliminary results on Tiny-ImageNet validate our approach's scalability to more complex datasets, offering a practical solution for developing robust image classification models.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Causal Interpretability for Adversarial Robustness: A Hybrid Generative Classification Approach
Zhao, Chunheng
Pisu, Pierluigi
Comert, Gurcan
Begashaw, Negash
Vaidyan, Varghese
Hubig, Nina Christine
Computer Vision and Pattern Recognition
Deep learning-based discriminative classifiers, despite their remarkable success, remain vulnerable to adversarial examples that can mislead model predictions. While adversarial training can enhance robustness, it fails to address the intrinsic vulnerability stemming from the opaque nature of these black-box models. We present a deep ensemble model that combines discriminative features with generative models to achieve both high accuracy and adversarial robustness. Our approach integrates a bottom-level pre-trained discriminative network for feature extraction with a top-level generative classification network that models adversarial input distributions through a deep latent variable model. Using variational Bayes, our model achieves superior robustness against white-box adversarial attacks without adversarial training. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate our model's superior adversarial robustness. Through evaluations using counterfactual metrics and feature interaction-based metrics, we establish correlations between model interpretability and adversarial robustness. Additionally, preliminary results on Tiny-ImageNet validate our approach's scalability to more complex datasets, offering a practical solution for developing robust image classification models.
title Causal Interpretability for Adversarial Robustness: A Hybrid Generative Classification Approach
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.20025