Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shang, Bingqi, Chen, Yiwei, Zhang, Yihua, Shen, Bingquan, Liu, Sijia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915563292327936
author Shang, Bingqi
Chen, Yiwei
Zhang, Yihua
Shen, Bingquan
Liu, Sijia
author_facet Shang, Bingqi
Chen, Yiwei
Zhang, Yihua
Shen, Bingquan
Liu, Sijia
contents Large language model (LLM) unlearning has become a critical mechanism for removing undesired data, knowledge, or behaviors from pre-trained models while retaining their general utility. Yet, with the rise of open-weight LLMs, we ask: can the unlearning process itself be backdoored, appearing successful under normal conditions yet reverting to pre-unlearned behavior when a hidden trigger is activated? Drawing inspiration from classical backdoor attacks that embed triggers into training data to enforce specific behaviors, we investigate backdoor unlearning, where models forget as intended in the clean setting but recover forgotten knowledge when the trigger appears. We show that designing such attacks presents unique challenges, hinging on where triggers are placed and how backdoor training is reinforced. We uncover a strong link between backdoor efficacy and the attention sink phenomenon, i.e., shallow input tokens consistently attract disproportionate attention in LLMs. Our analysis reveals that these attention sinks serve as gateways for backdoor unlearning: placing triggers at sink positions and aligning their attention values markedly enhances backdoor persistence. Extensive experiments validate these findings, showing that attention-sink-guided backdoor unlearning reliably restores forgotten knowledge in the presence of backdoor triggers, while behaving indistinguishably from a normally unlearned model when triggers are absent. Code is available at https://github.com/OPTML-Group/Unlearn-Backdoor.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
Shang, Bingqi
Chen, Yiwei
Zhang, Yihua
Shen, Bingquan
Liu, Sijia
Machine Learning
Computation and Language
Large language model (LLM) unlearning has become a critical mechanism for removing undesired data, knowledge, or behaviors from pre-trained models while retaining their general utility. Yet, with the rise of open-weight LLMs, we ask: can the unlearning process itself be backdoored, appearing successful under normal conditions yet reverting to pre-unlearned behavior when a hidden trigger is activated? Drawing inspiration from classical backdoor attacks that embed triggers into training data to enforce specific behaviors, we investigate backdoor unlearning, where models forget as intended in the clean setting but recover forgotten knowledge when the trigger appears. We show that designing such attacks presents unique challenges, hinging on where triggers are placed and how backdoor training is reinforced. We uncover a strong link between backdoor efficacy and the attention sink phenomenon, i.e., shallow input tokens consistently attract disproportionate attention in LLMs. Our analysis reveals that these attention sinks serve as gateways for backdoor unlearning: placing triggers at sink positions and aligning their attention values markedly enhances backdoor persistence. Extensive experiments validate these findings, showing that attention-sink-guided backdoor unlearning reliably restores forgotten knowledge in the presence of backdoor triggers, while behaving indistinguishably from a normally unlearned model when triggers are absent. Code is available at https://github.com/OPTML-Group/Unlearn-Backdoor.
title Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2510.17021