Universal Backdoor Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schneider, Benjamin, Lukas, Nils, Kerschbaum, Florian
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917571465314304
author Schneider, Benjamin
Lukas, Nils
Kerschbaum, Florian
author_facet Schneider, Benjamin
Lukas, Nils
Kerschbaum, Florian
contents Web-scraped datasets are vulnerable to data poisoning, which can be used for backdooring deep image classifiers during training. Since training on large datasets is expensive, a model is trained once and re-used many times. Unlike adversarial examples, backdoor attacks often target specific classes rather than any class learned by the model. One might expect that targeting many classes through a naive composition of attacks vastly increases the number of poison samples. We show this is not necessarily true and more efficient, universal data poisoning attacks exist that allow controlling misclassifications from any source class into any target class with a small increase in poison samples. Our idea is to generate triggers with salient characteristics that the model can learn. The triggers we craft exploit a phenomenon we call inter-class poison transferability, where learning a trigger from one class makes the model more vulnerable to learning triggers for other classes. We demonstrate the effectiveness and robustness of our universal backdoor attacks by controlling models with up to 6,000 classes while poisoning only 0.15% of the training dataset. Our source code is available at https://github.com/Ben-Schneider-code/Universal-Backdoor-Attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2312_00157
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Universal Backdoor Attacks
Schneider, Benjamin
Lukas, Nils
Kerschbaum, Florian
Machine Learning
Cryptography and Security
Computer Vision and Pattern Recognition
Web-scraped datasets are vulnerable to data poisoning, which can be used for backdooring deep image classifiers during training. Since training on large datasets is expensive, a model is trained once and re-used many times. Unlike adversarial examples, backdoor attacks often target specific classes rather than any class learned by the model. One might expect that targeting many classes through a naive composition of attacks vastly increases the number of poison samples. We show this is not necessarily true and more efficient, universal data poisoning attacks exist that allow controlling misclassifications from any source class into any target class with a small increase in poison samples. Our idea is to generate triggers with salient characteristics that the model can learn. The triggers we craft exploit a phenomenon we call inter-class poison transferability, where learning a trigger from one class makes the model more vulnerable to learning triggers for other classes. We demonstrate the effectiveness and robustness of our universal backdoor attacks by controlling models with up to 6,000 classes while poisoning only 0.15% of the training dataset. Our source code is available at https://github.com/Ben-Schneider-code/Universal-Backdoor-Attacks.
title Universal Backdoor Attacks
topic Machine Learning
Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.00157