Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shang, Bingqi, Chen, Yiwei, Zhang, Yihua, Shen, Bingquan, Liu, Sijia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PrivUn: Unveiling Latent Ripple Effects and Shallow Forgetting in Privacy Unlearning
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2026)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2026)
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
von: Liu, Yezi, et al.
Veröffentlicht: (2025)
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
Challenging Forgets: Unveiling the Worst-Case Forget Sets in Machine Unlearning
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
von: Fan, Chongyu, et al.
Veröffentlicht: (2025)
von: Fan, Chongyu, et al.
Veröffentlicht: (2025)
SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning
von: Jia, Jinghan, et al.
Veröffentlicht: (2024)
von: Jia, Jinghan, et al.
Veröffentlicht: (2024)
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
von: Chen, Yiwei, et al.
Veröffentlicht: (2025)
von: Chen, Yiwei, et al.
Veröffentlicht: (2025)
Forget Vectors at Play: Universal Input Perturbations Driving Machine Unlearning in Image Classification
von: Sun, Changchang, et al.
Veröffentlicht: (2024)
von: Sun, Changchang, et al.
Veröffentlicht: (2024)
Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities
von: Dang, Tien, et al.
Veröffentlicht: (2026)
von: Dang, Tien, et al.
Veröffentlicht: (2026)
To Forget or Not? Towards Practical Knowledge Unlearning for Large Language Models
von: Tian, Bozhong, et al.
Veröffentlicht: (2024)
von: Tian, Bozhong, et al.
Veröffentlicht: (2024)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
von: Wei, Rongzhe, et al.
Veröffentlicht: (2025)
von: Wei, Rongzhe, et al.
Veröffentlicht: (2025)
Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution
von: Sadhu, Saisab, et al.
Veröffentlicht: (2026)
von: Sadhu, Saisab, et al.
Veröffentlicht: (2026)
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
von: Reisizadeh, Hadi, et al.
Veröffentlicht: (2025)
von: Reisizadeh, Hadi, et al.
Veröffentlicht: (2025)
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
von: Xu, Shizhou, et al.
Veröffentlicht: (2025)
von: Xu, Shizhou, et al.
Veröffentlicht: (2025)
Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens
von: Koh, Seunghee, et al.
Veröffentlicht: (2026)
von: Koh, Seunghee, et al.
Veröffentlicht: (2026)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
von: He, Mutian, et al.
Veröffentlicht: (2025)
von: He, Mutian, et al.
Veröffentlicht: (2025)
FIT to Forget: Robust Continual Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
Machine Unlearning for Streaming Forgetting
von: Shen, Shaofei, et al.
Veröffentlicht: (2025)
von: Shen, Shaofei, et al.
Veröffentlicht: (2025)
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
von: Chen, Howard, et al.
Veröffentlicht: (2025)
von: Chen, Howard, et al.
Veröffentlicht: (2025)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
von: Lee, Chungpa, et al.
Veröffentlicht: (2026)
von: Lee, Chungpa, et al.
Veröffentlicht: (2026)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
von: Jo, Hyun-rae, et al.
Veröffentlicht: (2024)
von: Jo, Hyun-rae, et al.
Veröffentlicht: (2024)
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
von: Fan, Chongyu, et al.
Veröffentlicht: (2025)
von: Fan, Chongyu, et al.
Veröffentlicht: (2025)
RULE: Reinforcement UnLEarning Achieves Forget-Retain Pareto Optimality
von: Zhang, Chenlong, et al.
Veröffentlicht: (2025)
von: Zhang, Chenlong, et al.
Veröffentlicht: (2025)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
von: Liu, Yibai, et al.
Veröffentlicht: (2025)
von: Liu, Yibai, et al.
Veröffentlicht: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
Sparse Attention across Multiple-context KV Cache
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
von: Conklin, Henry C., et al.
Veröffentlicht: (2026)
von: Conklin, Henry C., et al.
Veröffentlicht: (2026)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
von: Watts, Ishaan, et al.
Veröffentlicht: (2026)
von: Watts, Ishaan, et al.
Veröffentlicht: (2026)
Graceful Forgetting in Generative Language Models
von: Jiang, Chunyang, et al.
Veröffentlicht: (2025)
von: Jiang, Chunyang, et al.
Veröffentlicht: (2025)
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
von: Hu, Shengyuan, et al.
Veröffentlicht: (2025)
von: Hu, Shengyuan, et al.
Veröffentlicht: (2025)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
SLIM: Let LLM Learn More and Forget Less with Soft LoRA and Identity Mixture
von: Han, Jiayi, et al.
Veröffentlicht: (2024)
von: Han, Jiayi, et al.
Veröffentlicht: (2024)
Learning and Forgetting Unsafe Examples in Large Language Models
von: Zhao, Jiachen, et al.
Veröffentlicht: (2023)
von: Zhao, Jiachen, et al.
Veröffentlicht: (2023)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PrivUn: Unveiling Latent Ripple Effects and Shallow Forgetting in Privacy Unlearning
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2026) -
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
von: Liu, Yezi, et al.
Veröffentlicht: (2025) -
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024) -
Challenging Forgets: Unveiling the Worst-Case Forget Sets in Machine Unlearning
von: Fan, Chongyu, et al.
Veröffentlicht: (2024) -
Forgetting Transformer: Softmax Attention with a Forget Gate
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)