PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wei, Chen, Pin-Yu, Liu, Sijia, Wang, Ren
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909580791906304
author Li, Wei
Chen, Pin-Yu
Liu, Sijia
Wang, Ren
author_facet Li, Wei
Chen, Pin-Yu
Liu, Sijia
Wang, Ren
contents Deep neural networks are susceptible to backdoor attacks, where adversaries manipulate model predictions by inserting malicious samples into the training data. Currently, there is still a significant challenge in identifying suspicious training data to unveil potential backdoor samples. In this paper, we propose a novel method, Prediction Shift Backdoor Detection (PSBD), leveraging an uncertainty-based approach requiring minimal unlabeled clean validation data. PSBD is motivated by an intriguing Prediction Shift (PS) phenomenon, where poisoned models' predictions on clean data often shift away from true labels towards certain other labels with dropout applied during inference, while backdoor samples exhibit less PS. We hypothesize PS results from the neuron bias effect, making neurons favor features of certain classes. PSBD identifies backdoor training samples by computing the Prediction Shift Uncertainty (PSU), the variance in probability values when dropout layers are toggled on and off during model inference. Extensive experiments have been conducted to verify the effectiveness and efficiency of PSBD, which achieves state-of-the-art results among mainstream detection methods. The code is available at https://github.com/WL-619/PSBD.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05826
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
Li, Wei
Chen, Pin-Yu
Liu, Sijia
Wang, Ren
Machine Learning
Artificial Intelligence
Cryptography and Security
Deep neural networks are susceptible to backdoor attacks, where adversaries manipulate model predictions by inserting malicious samples into the training data. Currently, there is still a significant challenge in identifying suspicious training data to unveil potential backdoor samples. In this paper, we propose a novel method, Prediction Shift Backdoor Detection (PSBD), leveraging an uncertainty-based approach requiring minimal unlabeled clean validation data. PSBD is motivated by an intriguing Prediction Shift (PS) phenomenon, where poisoned models' predictions on clean data often shift away from true labels towards certain other labels with dropout applied during inference, while backdoor samples exhibit less PS. We hypothesize PS results from the neuron bias effect, making neurons favor features of certain classes. PSBD identifies backdoor training samples by computing the Prediction Shift Uncertainty (PSU), the variance in probability values when dropout layers are toggled on and off during model inference. Extensive experiments have been conducted to verify the effectiveness and efficiency of PSBD, which achieves state-of-the-art results among mainstream detection methods. The code is available at https://github.com/WL-619/PSBD.
title PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2406.05826