In-distribution adversarial attacks on object recognition models using gradient-free search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Madan, Spandan, Sasaki, Tomotake, Pfister, Hanspeter, Li, Tzu-Mao, Boix, Xavier
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912211412189184
author Madan, Spandan
Sasaki, Tomotake
Pfister, Hanspeter
Li, Tzu-Mao
Boix, Xavier
author_facet Madan, Spandan
Sasaki, Tomotake
Pfister, Hanspeter
Li, Tzu-Mao
Boix, Xavier
contents Neural networks are susceptible to small perturbations in the form of 2D rotations and shifts, image crops, and even changes in object colors. Past works attribute these errors to dataset bias, claiming that models fail on these perturbed samples as they do not belong to the training data distribution. Here, we challenge this claim and present evidence of the widespread existence of perturbed images within the training data distribution, which networks fail to classify. We train models on data sampled from parametric distributions, then search inside this data distribution to find such in-distribution adversarial examples. This is done using our gradient-free evolution strategies (ES) based approach which we call CMA-Search. Despite training with a large-scale (0.5 million images), unbiased dataset of camera and light variations, CMA-Search can find a failure inside the data distribution in over 71% cases by perturbing the camera position. With lighting changes, CMA-Search finds misclassifications in 42% cases. These findings also extend to natural images from ImageNet and Co3D datasets. This phenomenon of in-distribution images presents a highly worrisome problem for artificial intelligence -- they bypass the need for a malicious agent to add engineered noise to induce an adversarial attack. All code, datasets, and demos are available at https://github.com/Spandan-Madan/in_distribution_adversarial_examples.
format Preprint
id arxiv_https___arxiv_org_abs_2106_16198
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle In-distribution adversarial attacks on object recognition models using gradient-free search
Madan, Spandan
Sasaki, Tomotake
Pfister, Hanspeter
Li, Tzu-Mao
Boix, Xavier
Computer Vision and Pattern Recognition
Machine Learning
Neural networks are susceptible to small perturbations in the form of 2D rotations and shifts, image crops, and even changes in object colors. Past works attribute these errors to dataset bias, claiming that models fail on these perturbed samples as they do not belong to the training data distribution. Here, we challenge this claim and present evidence of the widespread existence of perturbed images within the training data distribution, which networks fail to classify. We train models on data sampled from parametric distributions, then search inside this data distribution to find such in-distribution adversarial examples. This is done using our gradient-free evolution strategies (ES) based approach which we call CMA-Search. Despite training with a large-scale (0.5 million images), unbiased dataset of camera and light variations, CMA-Search can find a failure inside the data distribution in over 71% cases by perturbing the camera position. With lighting changes, CMA-Search finds misclassifications in 42% cases. These findings also extend to natural images from ImageNet and Co3D datasets. This phenomenon of in-distribution images presents a highly worrisome problem for artificial intelligence -- they bypass the need for a malicious agent to add engineered noise to induce an adversarial attack. All code, datasets, and demos are available at https://github.com/Spandan-Madan/in_distribution_adversarial_examples.
title In-distribution adversarial attacks on object recognition models using gradient-free search
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2106.16198