From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Siddiqui, Shoaib Ahmed, Weller, Adrian, Krueger, David, Dziugaite, Gintare Karolina, Mozer, Michael Curtis, Triantafillou, Eleni
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915730450022400
author Siddiqui, Shoaib Ahmed
Weller, Adrian
Krueger, David
Dziugaite, Gintare Karolina
Mozer, Michael Curtis
Triantafillou, Eleni
author_facet Siddiqui, Shoaib Ahmed
Weller, Adrian
Krueger, David
Dziugaite, Gintare Karolina
Mozer, Michael Curtis
Triantafillou, Eleni
contents Recent unlearning methods for LLMs are vulnerable to relearning attacks: knowledge believed-to-be-unlearned re-emerges by fine-tuning on a small set of (even seemingly-unrelated) examples. We study this phenomenon in a controlled setting for example-level unlearning in vision classifiers. We make the surprising discovery that forget-set accuracy can recover from around 50% post-unlearning to nearly 100% with fine-tuning on just the retain set -- i.e., zero examples of the forget set. We observe this effect across a wide variety of unlearning methods, whereas for a model retrained from scratch excluding the forget set (gold standard), the accuracy remains at 50%. We observe that resistance to relearning attacks can be predicted by weight-space properties, specifically, $L_2$-distance and linear mode connectivity between the original and the unlearned model. Leveraging this insight, we propose a new class of methods that achieve state-of-the-art resistance to relearning attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22310
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
Siddiqui, Shoaib Ahmed
Weller, Adrian
Krueger, David
Dziugaite, Gintare Karolina
Mozer, Michael Curtis
Triantafillou, Eleni
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Recent unlearning methods for LLMs are vulnerable to relearning attacks: knowledge believed-to-be-unlearned re-emerges by fine-tuning on a small set of (even seemingly-unrelated) examples. We study this phenomenon in a controlled setting for example-level unlearning in vision classifiers. We make the surprising discovery that forget-set accuracy can recover from around 50% post-unlearning to nearly 100% with fine-tuning on just the retain set -- i.e., zero examples of the forget set. We observe this effect across a wide variety of unlearning methods, whereas for a model retrained from scratch excluding the forget set (gold standard), the accuracy remains at 50%. We observe that resistance to relearning attacks can be predicted by weight-space properties, specifically, $L_2$-distance and linear mode connectivity between the original and the unlearned model. Leveraging this insight, we propose a new class of methods that achieve state-of-the-art resistance to relearning attacks.
title From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22310