Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Changsheng, Zhang, Yihua, Jia, Jinghan, Ram, Parikshit, Wei, Dennis, Yao, Yuguang, Pal, Soumyadeep, Baracaldo, Nathalie, Liu, Sijia
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911203337437184
author Wang, Changsheng
Zhang, Yihua
Jia, Jinghan
Ram, Parikshit
Wei, Dennis
Yao, Yuguang
Pal, Soumyadeep
Baracaldo, Nathalie
Liu, Sijia
author_facet Wang, Changsheng
Zhang, Yihua
Jia, Jinghan
Ram, Parikshit
Wei, Dennis
Yao, Yuguang
Pal, Soumyadeep
Baracaldo, Nathalie
Liu, Sijia
contents Machine unlearning offers a promising solution to privacy and safety concerns in large language models (LLMs) by selectively removing targeted knowledge while preserving utility. However, current methods are highly sensitive to downstream fine-tuning, which can quickly recover forgotten information-even from unrelated tasks. To address this, we introduce invariance into unlearning for the first time, inspired by invariant risk minimization (IRM). Building on this principle, we propose invariant LLM unlearning (ILU), a regularization-based framework that enhances robustness. Notably, ILU generalizes well to diverse fine-tuning tasks, even when trained using a single dataset. A task vector analysis is also provided to further elucidate the rationale behind ILU's effectiveness. Extensive experiments on the WMDP and MUSE benchmark, reveal that ILU significantly outperforms state-of-the-art unlearning methods, including negative preference optimization (NPO) and representation misdirection for unlearning (RMU). Notably, ILU achieves superior unlearning robustness across diverse downstream fine-tuning scenarios (e.g., math, paraphrase detection, and sentiment analysis) while preserving the fine-tuning performance.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
Wang, Changsheng
Zhang, Yihua
Jia, Jinghan
Ram, Parikshit
Wei, Dennis
Yao, Yuguang
Pal, Soumyadeep
Baracaldo, Nathalie
Liu, Sijia
Machine Learning
Machine unlearning offers a promising solution to privacy and safety concerns in large language models (LLMs) by selectively removing targeted knowledge while preserving utility. However, current methods are highly sensitive to downstream fine-tuning, which can quickly recover forgotten information-even from unrelated tasks. To address this, we introduce invariance into unlearning for the first time, inspired by invariant risk minimization (IRM). Building on this principle, we propose invariant LLM unlearning (ILU), a regularization-based framework that enhances robustness. Notably, ILU generalizes well to diverse fine-tuning tasks, even when trained using a single dataset. A task vector analysis is also provided to further elucidate the rationale behind ILU's effectiveness. Extensive experiments on the WMDP and MUSE benchmark, reveal that ILU significantly outperforms state-of-the-art unlearning methods, including negative preference optimization (NPO) and representation misdirection for unlearning (RMU). Notably, ILU achieves superior unlearning robustness across diverse downstream fine-tuning scenarios (e.g., math, paraphrase detection, and sentiment analysis) while preserving the fine-tuning performance.
title Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
topic Machine Learning
url https://arxiv.org/abs/2506.01339