Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jiang, Peihai, Lyu, Xixiang, Li, Yige, Ma, Jing
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929661375676416
author Jiang, Peihai
Lyu, Xixiang
Li, Yige
Ma, Jing
author_facet Jiang, Peihai
Lyu, Xixiang
Li, Yige
Ma, Jing
contents Supervised fine-tuning has become the predominant method for adapting large pretrained models to downstream tasks. However, recent studies have revealed that these models are vulnerable to backdoor attacks, where even a small number of malicious samples can successfully embed backdoor triggers into the model. While most existing defense methods focus on post-training backdoor defense, efficiently defending against backdoor attacks during training phase remains largely unexplored. To address this gap, we propose a novel defense method called Backdoor Token Unlearning (BTU), which proactively detects and neutralizes trigger tokens during the training stage. Our work is based on two key findings: 1) backdoor learning causes distinctive differences between backdoor token parameters and clean token parameters in word embedding layers, and 2) the success of backdoor attacks heavily depends on backdoor token parameters. The BTU defense leverages these properties to identify aberrant embedding parameters and subsequently removes backdoor behaviors using a fine-grained unlearning technique. Extensive evaluations across three datasets and four types of backdoor attacks demonstrate that BTU effectively defends against these threats while preserving the model's performance on primary tasks. Our code is available at https://github.com/XDJPH/BTU.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03272
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
Jiang, Peihai
Lyu, Xixiang
Li, Yige
Ma, Jing
Cryptography and Security
Artificial Intelligence
Computation and Language
Supervised fine-tuning has become the predominant method for adapting large pretrained models to downstream tasks. However, recent studies have revealed that these models are vulnerable to backdoor attacks, where even a small number of malicious samples can successfully embed backdoor triggers into the model. While most existing defense methods focus on post-training backdoor defense, efficiently defending against backdoor attacks during training phase remains largely unexplored. To address this gap, we propose a novel defense method called Backdoor Token Unlearning (BTU), which proactively detects and neutralizes trigger tokens during the training stage. Our work is based on two key findings: 1) backdoor learning causes distinctive differences between backdoor token parameters and clean token parameters in word embedding layers, and 2) the success of backdoor attacks heavily depends on backdoor token parameters. The BTU defense leverages these properties to identify aberrant embedding parameters and subsequently removes backdoor behaviors using a fine-grained unlearning technique. Extensive evaluations across three datasets and four types of backdoor attacks demonstrate that BTU effectively defends against these threats while preserving the model's performance on primary tasks. Our code is available at https://github.com/XDJPH/BTU.
title Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2501.03272