Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jie, Schlarmann, Christian, Nikolić, Kristina, Carlini, Nicholas, Croce, Francesco, Hein, Matthias, Tramèr, Florian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917910879928320
author Zhang, Jie
Schlarmann, Christian
Nikolić, Kristina
Carlini, Nicholas
Croce, Francesco
Hein, Matthias
Tramèr, Florian
author_facet Zhang, Jie
Schlarmann, Christian
Nikolić, Kristina
Carlini, Nicholas
Croce, Francesco
Hein, Matthias
Tramèr, Florian
contents Ensemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations at multiple noisy image resolutions, producing a single robust classification. This defense was shown to be effective against multiple state-of-the-art attacks. Perhaps even more convincingly, it was shown that the model's gradients are perceptually aligned: attacks against the model produce noise that perceptually resembles the targeted class. In this short note, we show that this defense is not robust to adversarial attack. We first show that the defense's randomness and ensembling method cause severe gradient masking. We then use standard adaptive attack techniques to reduce the defense's robust accuracy from 48% to 14% on CIFAR-100 and from 62% to 11% on CIFAR-10, under the $\ell_\infty$-norm threat model with $\varepsilon=8/255$.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14834
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
Zhang, Jie
Schlarmann, Christian
Nikolić, Kristina
Carlini, Nicholas
Croce, Francesco
Hein, Matthias
Tramèr, Florian
Machine Learning
Cryptography and Security
Ensemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermediate representations at multiple noisy image resolutions, producing a single robust classification. This defense was shown to be effective against multiple state-of-the-art attacks. Perhaps even more convincingly, it was shown that the model's gradients are perceptually aligned: attacks against the model produce noise that perceptually resembles the targeted class. In this short note, we show that this defense is not robust to adversarial attack. We first show that the defense's randomness and ensembling method cause severe gradient masking. We then use standard adaptive attack techniques to reduce the defense's robust accuracy from 48% to 14% on CIFAR-100 and from 62% to 11% on CIFAR-10, under the $\ell_\infty$-norm threat model with $\varepsilon=8/255$.
title Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2411.14834