Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Xiaoyu, Yue, Xiang, Liu, Yang, Ye, Qingqing, Zheng, Huadi, Hu, Peizhao, Du, Minxin, Hu, Haibo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917501253713920
author Xu, Xiaoyu
Yue, Xiang
Liu, Yang
Ye, Qingqing
Zheng, Huadi
Hu, Peizhao
Du, Minxin
Hu, Haibo
author_facet Xu, Xiaoyu
Yue, Xiang
Liu, Yang
Ye, Qingqing
Zheng, Huadi
Hu, Peizhao
Du, Minxin
Hu, Haibo
contents Unlearning in large language models (LLMs) aims to remove specified data, but its efficacy is typically assessed with task-level metrics like accuracy and perplexity. We show that these metrics can be misleading, as models can appear to forget while their original behavior is easily restored through minimal fine-tuning. This \emph{reversibility} suggests that information is merely suppressed, not genuinely erased. To address this critical evaluation gap, we introduce a \emph{representation-level analysis framework}. Our toolkit comprises PCA similarity and shift, centered kernel alignment (CKA), and Fisher information, complemented by a summary metric, the mean PCA distance, to measure representational drift. Applying this framework across multiple unlearning methods, data domains, and LLMs, we identify four distinct forgetting regimes based on their \emph{reversibility} and \emph{catastrophicity}. We compare recovery strategies and show that relearning efficiency relies on the data source. We also find that irreversible, non-catastrophic forgetting is exceptionally challenging. By probing unlearning limits, we identify a case of seemingly irreversible, targeted forgetting, offering insights for more robust erasure algorithms. Overall, our findings expose a gap in current evaluation and establish a representation-level foundation for trustworthy unlearning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16831
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
Xu, Xiaoyu
Yue, Xiang
Liu, Yang
Ye, Qingqing
Zheng, Huadi
Hu, Peizhao
Du, Minxin
Hu, Haibo
Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
Unlearning in large language models (LLMs) aims to remove specified data, but its efficacy is typically assessed with task-level metrics like accuracy and perplexity. We show that these metrics can be misleading, as models can appear to forget while their original behavior is easily restored through minimal fine-tuning. This \emph{reversibility} suggests that information is merely suppressed, not genuinely erased. To address this critical evaluation gap, we introduce a \emph{representation-level analysis framework}. Our toolkit comprises PCA similarity and shift, centered kernel alignment (CKA), and Fisher information, complemented by a summary metric, the mean PCA distance, to measure representational drift. Applying this framework across multiple unlearning methods, data domains, and LLMs, we identify four distinct forgetting regimes based on their \emph{reversibility} and \emph{catastrophicity}. We compare recovery strategies and show that relearning efficiency relies on the data source. We also find that irreversible, non-catastrophic forgetting is exceptionally challenging. By probing unlearning limits, we identify a case of seemingly irreversible, targeted forgetting, offering insights for more robust erasure algorithms. Overall, our findings expose a gap in current evaluation and establish a representation-level foundation for trustworthy unlearning.
title Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
topic Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2505.16831