Prototype-Guided Robust Learning against Backdoor Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Wei, Pintor, Maura, Demontis, Ambra, Biggio, Battista
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915956792492032
author Guo, Wei
Pintor, Maura
Demontis, Ambra
Biggio, Battista
author_facet Guo, Wei
Pintor, Maura
Demontis, Ambra
Biggio, Battista
contents Backdoor attacks poison the training data, causing the model to behave normally on clean inputs but predict attacker-chosen labels when trigger patterns are embedded into the input samples. Defending against such attacks is highly challenging, especially when the defender has limited access to clean data. Existing defense methods often rely on restrictive assumptions-such as high poisoning ratios or poisoning strategies-limiting their practicality and generalization. To overcome these limitations, we propose Prototype-Guided Robust Learning (PGRL), a defense that only requires a small set of verified benign samples, and integrates two complementary components during fine-tuning: Label Consistency Verification (LCV), which detects and removes suspicious samples from the potentially poisoned dataset; and Feature Distance Estimation (FDE), which enforces the unlearning of backdoor-related representations. Extensive experiments against eight existing defenses show that PGRL achieves superior robustness across diverse architectures, datasets, and advanced attack scenarios, establishing a new standard for practical and generalizable backdoor defense.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08748
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prototype-Guided Robust Learning against Backdoor Attacks
Guo, Wei
Pintor, Maura
Demontis, Ambra
Biggio, Battista
Cryptography and Security
Backdoor attacks poison the training data, causing the model to behave normally on clean inputs but predict attacker-chosen labels when trigger patterns are embedded into the input samples. Defending against such attacks is highly challenging, especially when the defender has limited access to clean data. Existing defense methods often rely on restrictive assumptions-such as high poisoning ratios or poisoning strategies-limiting their practicality and generalization. To overcome these limitations, we propose Prototype-Guided Robust Learning (PGRL), a defense that only requires a small set of verified benign samples, and integrates two complementary components during fine-tuning: Label Consistency Verification (LCV), which detects and removes suspicious samples from the potentially poisoned dataset; and Feature Distance Estimation (FDE), which enforces the unlearning of backdoor-related representations. Extensive experiments against eight existing defenses show that PGRL achieves superior robustness across diverse architectures, datasets, and advanced attack scenarios, establishing a new standard for practical and generalizable backdoor defense.
title Prototype-Guided Robust Learning against Backdoor Attacks
topic Cryptography and Security
url https://arxiv.org/abs/2509.08748