Prototype Guided Backdoor Defense

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Amula, Venkat Adithya, Samavedam, Sunayana, Saini, Saurabh, Gupta, Avani, J, Narayanan P
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908285812080640
author Amula, Venkat Adithya
Samavedam, Sunayana
Saini, Saurabh
Gupta, Avani
J, Narayanan P
author_facet Amula, Venkat Adithya
Samavedam, Sunayana
Saini, Saurabh
Gupta, Avani
J, Narayanan P
contents Deep learning models are susceptible to {\em backdoor attacks} involving malicious attackers perturbing a small subset of training data with a {\em trigger} to causes misclassifications. Various triggers have been used, including semantic triggers that are easily realizable without requiring the attacker to manipulate the image. The emergence of generative AI has eased the generation of varied poisoned samples. Robustness across types of triggers is crucial to effective defense. We propose Prototype Guided Backdoor Defense (PGBD), a robust post-hoc defense that scales across different trigger types, including previously unsolved semantic triggers. PGBD exploits displacements in the geometric spaces of activations to penalize movements toward the trigger. This is done using a novel sanitization loss of a post-hoc fine-tuning step. The geometric approach scales easily to all types of attacks. PGBD achieves better performance across all settings. We also present the first defense against a new semantic attack on celebrity face images. Project page: \hyperlink{https://venkatadithya9.github.io/pgbd.github.io/}{this https URL}.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20925
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prototype Guided Backdoor Defense
Amula, Venkat Adithya
Samavedam, Sunayana
Saini, Saurabh
Gupta, Avani
J, Narayanan P
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Deep learning models are susceptible to {\em backdoor attacks} involving malicious attackers perturbing a small subset of training data with a {\em trigger} to causes misclassifications. Various triggers have been used, including semantic triggers that are easily realizable without requiring the attacker to manipulate the image. The emergence of generative AI has eased the generation of varied poisoned samples. Robustness across types of triggers is crucial to effective defense. We propose Prototype Guided Backdoor Defense (PGBD), a robust post-hoc defense that scales across different trigger types, including previously unsolved semantic triggers. PGBD exploits displacements in the geometric spaces of activations to penalize movements toward the trigger. This is done using a novel sanitization loss of a post-hoc fine-tuning step. The geometric approach scales easily to all types of attacks. PGBD achieves better performance across all settings. We also present the first defense against a new semantic attack on celebrity face images. Project page: \hyperlink{https://venkatadithya9.github.io/pgbd.github.io/}{this https URL}.
title Prototype Guided Backdoor Defense
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.20925