Adversarial Doodles: Interpretable and Human-drawable Attacks Provide Describable Insights

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nara, Ryoya, Matsui, Yusuke
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929495431184384
author Nara, Ryoya
Matsui, Yusuke
author_facet Nara, Ryoya
Matsui, Yusuke
contents DNN-based image classifiers are susceptible to adversarial attacks. Most previous adversarial attacks do not have clear patterns, making it difficult to interpret attacks' results and gain insights into classifiers' mechanisms. Therefore, we propose Adversarial Doodles, which have interpretable shapes. We optimize black bezier curves to fool the classifier by overlaying them onto the input image. By introducing random affine transformation and regularizing the doodled area, we obtain small-sized attacks that cause misclassification even when humans replicate them by hand. Adversarial doodles provide describable insights into the relationship between the human-drawn doodle's shape and the classifier's output, such as "When we add three small circles on a helicopter image, the ResNet-50 classifier mistakenly classifies it as an airplane."
format Preprint
id arxiv_https___arxiv_org_abs_2311_15994
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Adversarial Doodles: Interpretable and Human-drawable Attacks Provide Describable Insights
Nara, Ryoya
Matsui, Yusuke
Computer Vision and Pattern Recognition
DNN-based image classifiers are susceptible to adversarial attacks. Most previous adversarial attacks do not have clear patterns, making it difficult to interpret attacks' results and gain insights into classifiers' mechanisms. Therefore, we propose Adversarial Doodles, which have interpretable shapes. We optimize black bezier curves to fool the classifier by overlaying them onto the input image. By introducing random affine transformation and regularizing the doodled area, we obtain small-sized attacks that cause misclassification even when humans replicate them by hand. Adversarial doodles provide describable insights into the relationship between the human-drawn doodle's shape and the classifier's output, such as "When we add three small circles on a helicopter image, the ResNet-50 classifier mistakenly classifies it as an airplane."
title Adversarial Doodles: Interpretable and Human-drawable Attacks Provide Describable Insights
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.15994