Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Shengyuan, Kale, Neil, Thaker, Pratiksha, Fu, Yiwei, Wu, Steven, Smith, Virginia
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2506.15699
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913900800245760
author Hu, Shengyuan
Kale, Neil
Thaker, Pratiksha
Fu, Yiwei
Wu, Steven
Smith, Virginia
author_facet Hu, Shengyuan
Kale, Neil
Thaker, Pratiksha
Fu, Yiwei
Wu, Steven
Smith, Virginia
contents Machine unlearning has the potential to improve the safety of large language models (LLMs) by removing sensitive or harmful information post hoc. A key challenge in unlearning involves balancing between forget quality (effectively unlearning undesirable information) and retain quality (maintaining good performance on other, general tasks). Unfortunately, as we show, current LLM unlearning benchmarks contain highly disparate forget and retain sets -- painting a false picture of the effectiveness of LLM unlearning methods. This can be particularly problematic because it opens the door for benign perturbations, such as relearning attacks, to easily reveal supposedly unlearned knowledge once models are deployed. To address this, we present $\texttt{BLUR}$: a benchmark for LLM unlearning that provides more realistic scenarios of forget-retain overlap. $\texttt{BLUR}$ significantly expands on existing unlearning benchmarks by providing extended evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on $\texttt{BLUR}$, with simple approaches performing better on average than more recent methods. These results highlight the importance of robust evaluation and suggest several important directions of future study. Our benchmark is publicly available at: https://huggingface.co/datasets/forgelab/BLUR
format Preprint
id arxiv_https___arxiv_org_abs_2506_15699
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
Hu, Shengyuan
Kale, Neil
Thaker, Pratiksha
Fu, Yiwei
Wu, Steven
Smith, Virginia
Machine Learning
Artificial Intelligence
Machine unlearning has the potential to improve the safety of large language models (LLMs) by removing sensitive or harmful information post hoc. A key challenge in unlearning involves balancing between forget quality (effectively unlearning undesirable information) and retain quality (maintaining good performance on other, general tasks). Unfortunately, as we show, current LLM unlearning benchmarks contain highly disparate forget and retain sets -- painting a false picture of the effectiveness of LLM unlearning methods. This can be particularly problematic because it opens the door for benign perturbations, such as relearning attacks, to easily reveal supposedly unlearned knowledge once models are deployed. To address this, we present $\texttt{BLUR}$: a benchmark for LLM unlearning that provides more realistic scenarios of forget-retain overlap. $\texttt{BLUR}$ significantly expands on existing unlearning benchmarks by providing extended evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on $\texttt{BLUR}$, with simple approaches performing better on average than more recent methods. These results highlight the importance of robust evaluation and suggest several important directions of future study. Our benchmark is publicly available at: https://huggingface.co/datasets/forgelab/BLUR
title BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.15699