Existing Large Language Model Unlearning Evaluations Are Inconclusive

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Zhili, Xu, Yixuan Even, Robey, Alexander, Kirk, Robert, Davies, Xander, Gal, Yarin, Schwarzschild, Avi, Kolter, J. Zico
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908388053483520
author Feng, Zhili
Xu, Yixuan Even
Robey, Alexander
Kirk, Robert
Davies, Xander
Gal, Yarin
Schwarzschild, Avi
Kolter, J. Zico
author_facet Feng, Zhili
Xu, Yixuan Even
Robey, Alexander
Kirk, Robert
Davies, Xander
Gal, Yarin
Schwarzschild, Avi
Kolter, J. Zico
contents Machine unlearning aims to remove sensitive or undesired data from large language models. However, recent studies suggest that unlearning is often shallow, claiming that removed knowledge can easily be recovered. In this work, we critically examine standard unlearning evaluation practices and uncover key limitations that shake our trust in those findings. First, we show that some evaluations introduce substantial new information into the model, potentially masking true unlearning performance by re-teaching the model during testing. Second, we demonstrate that evaluation outcomes vary significantly across tasks, undermining the generalizability of current evaluation routines. Finally, we find that many evaluations rely on spurious correlations, making their results difficult to trust and interpret. Taken together, these issues suggest that current evaluation protocols may both overstate and understate unlearning success. To address this, we propose two principles for future unlearning evaluations: minimal information injection and downstream task awareness. We validate these principles through a series of targeted experiments, showing how violations of each can lead to misleading conclusions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00688
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Existing Large Language Model Unlearning Evaluations Are Inconclusive
Feng, Zhili
Xu, Yixuan Even
Robey, Alexander
Kirk, Robert
Davies, Xander
Gal, Yarin
Schwarzschild, Avi
Kolter, J. Zico
Machine Learning
Artificial Intelligence
Computation and Language
Machine unlearning aims to remove sensitive or undesired data from large language models. However, recent studies suggest that unlearning is often shallow, claiming that removed knowledge can easily be recovered. In this work, we critically examine standard unlearning evaluation practices and uncover key limitations that shake our trust in those findings. First, we show that some evaluations introduce substantial new information into the model, potentially masking true unlearning performance by re-teaching the model during testing. Second, we demonstrate that evaluation outcomes vary significantly across tasks, undermining the generalizability of current evaluation routines. Finally, we find that many evaluations rely on spurious correlations, making their results difficult to trust and interpret. Taken together, these issues suggest that current evaluation protocols may both overstate and understate unlearning success. To address this, we propose two principles for future unlearning evaluations: minimal information injection and downstream task awareness. We validate these principles through a series of targeted experiments, showing how violations of each can lead to misleading conclusions.
title Existing Large Language Model Unlearning Evaluations Are Inconclusive
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.00688