Backdoor Learning Curves: Explaining Backdoor Poisoning Beyond Influence Functions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cinà, Antonio Emanuele, Grosse, Kathrin, Vascon, Sebastiano, Demontis, Ambra, Biggio, Battista, Roli, Fabio, Pelillo, Marcello
Natura: Preprint
Pubblicazione: 2021
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916523741806592
author Cinà, Antonio Emanuele
Grosse, Kathrin
Vascon, Sebastiano
Demontis, Ambra
Biggio, Battista
Roli, Fabio
Pelillo, Marcello
author_facet Cinà, Antonio Emanuele
Grosse, Kathrin
Vascon, Sebastiano
Demontis, Ambra
Biggio, Battista
Roli, Fabio
Pelillo, Marcello
contents Backdoor attacks inject poisoning samples during training, with the goal of forcing a machine learning model to output an attacker-chosen class when presented a specific trigger at test time. Although backdoor attacks have been demonstrated in a variety of settings and against different models, the factors affecting their effectiveness are still not well understood. In this work, we provide a unifying framework to study the process of backdoor learning under the lens of incremental learning and influence functions. We show that the effectiveness of backdoor attacks depends on: (i) the complexity of the learning algorithm, controlled by its hyperparameters; (ii) the fraction of backdoor samples injected into the training set; and (iii) the size and visibility of the backdoor trigger. These factors affect how fast a model learns to correlate the presence of the backdoor trigger with the target class. Our analysis unveils the intriguing existence of a region in the hyperparameter space in which the accuracy on clean test samples is still high while backdoor attacks are ineffective, thereby suggesting novel criteria to improve existing defenses.
format Preprint
id arxiv_https___arxiv_org_abs_2106_07214
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Backdoor Learning Curves: Explaining Backdoor Poisoning Beyond Influence Functions
Cinà, Antonio Emanuele
Grosse, Kathrin
Vascon, Sebastiano
Demontis, Ambra
Biggio, Battista
Roli, Fabio
Pelillo, Marcello
Machine Learning
Cryptography and Security
Backdoor attacks inject poisoning samples during training, with the goal of forcing a machine learning model to output an attacker-chosen class when presented a specific trigger at test time. Although backdoor attacks have been demonstrated in a variety of settings and against different models, the factors affecting their effectiveness are still not well understood. In this work, we provide a unifying framework to study the process of backdoor learning under the lens of incremental learning and influence functions. We show that the effectiveness of backdoor attacks depends on: (i) the complexity of the learning algorithm, controlled by its hyperparameters; (ii) the fraction of backdoor samples injected into the training set; and (iii) the size and visibility of the backdoor trigger. These factors affect how fast a model learns to correlate the presence of the backdoor trigger with the target class. Our analysis unveils the intriguing existence of a region in the hyperparameter space in which the accuracy on clean test samples is still high while backdoor attacks are ineffective, thereby suggesting novel criteria to improve existing defenses.
title Backdoor Learning Curves: Explaining Backdoor Poisoning Beyond Influence Functions
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2106.07214