Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Song, Baogang, Zhao, Dongdong, Xiang, Jianwen, Xu, Qiben, Yu, Zizhuo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914094106279936
author Song, Baogang
Zhao, Dongdong
Xiang, Jianwen
Xu, Qiben
Yu, Zizhuo
author_facet Song, Baogang
Zhao, Dongdong
Xiang, Jianwen
Xu, Qiben
Yu, Zizhuo
contents Backdoor attacks pose a persistent security risk to deep neural networks (DNNs) due to their stealth and durability. While recent research has explored leveraging model unlearning mechanisms to enhance backdoor concealment, existing attack strategies still leave persistent traces that may be detected through static analysis. In this work, we introduce the first paradigm of revocable backdoor attacks, where the backdoor can be proactively and thoroughly removed after the attack objective is achieved. We formulate the trigger optimization in revocable backdoor attacks as a bilevel optimization problem: by simulating both backdoor injection and unlearning processes, the trigger generator is optimized to achieve a high attack success rate (ASR) while ensuring that the backdoor can be easily erased through unlearning. To mitigate the optimization conflict between injection and removal objectives, we employ a deterministic partition of poisoning and unlearning samples to reduce sampling-induced variance, and further apply the Projected Conflicting Gradient (PCGrad) technique to resolve the remaining gradient conflicts. Experiments on CIFAR-10 and ImageNet demonstrate that our method maintains ASR comparable to state-of-the-art backdoor attacks, while enabling effective removal of backdoor behavior after unlearning. This work opens a new direction for backdoor attack research and presents new challenges for the security of machine learning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13322
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
Song, Baogang
Zhao, Dongdong
Xiang, Jianwen
Xu, Qiben
Yu, Zizhuo
Cryptography and Security
Artificial Intelligence
Backdoor attacks pose a persistent security risk to deep neural networks (DNNs) due to their stealth and durability. While recent research has explored leveraging model unlearning mechanisms to enhance backdoor concealment, existing attack strategies still leave persistent traces that may be detected through static analysis. In this work, we introduce the first paradigm of revocable backdoor attacks, where the backdoor can be proactively and thoroughly removed after the attack objective is achieved. We formulate the trigger optimization in revocable backdoor attacks as a bilevel optimization problem: by simulating both backdoor injection and unlearning processes, the trigger generator is optimized to achieve a high attack success rate (ASR) while ensuring that the backdoor can be easily erased through unlearning. To mitigate the optimization conflict between injection and removal objectives, we employ a deterministic partition of poisoning and unlearning samples to reduce sampling-induced variance, and further apply the Projected Conflicting Gradient (PCGrad) technique to resolve the remaining gradient conflicts. Experiments on CIFAR-10 and ImageNet demonstrate that our method maintains ASR comparable to state-of-the-art backdoor attacks, while enabling effective removal of backdoor behavior after unlearning. This work opens a new direction for backdoor attack research and presents new challenges for the security of machine learning systems.
title Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.13322