OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dorna, Vineeth, Mekala, Anmol, Zhao, Wenlong, McCallum, Andrew, Lipton, Zachary C., Kolter, J. Zico, Maini, Pratyush
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917070094991360
author Dorna, Vineeth
Mekala, Anmol
Zhao, Wenlong
McCallum, Andrew
Lipton, Zachary C.
Kolter, J. Zico
Maini, Pratyush
author_facet Dorna, Vineeth
Mekala, Anmol
Zhao, Wenlong
McCallum, Andrew
Lipton, Zachary C.
Kolter, J. Zico
Maini, Pratyush
contents Robust unlearning is crucial for safely deploying large language models (LLMs) in environments where data privacy, model safety, and regulatory compliance must be ensured. Yet the task is inherently challenging, partly due to difficulties in reliably measuring whether unlearning has truly occurred. Moreover, fragmentation in current methodologies and inconsistent evaluation metrics hinder comparative analysis and reproducibility. To unify and accelerate research efforts, we introduce OpenUnlearning, a standardized and extensible framework designed explicitly for benchmarking both LLM unlearning methods and metrics. OpenUnlearning integrates 13 unlearning algorithms and 16 diverse evaluations across 3 leading benchmarks (TOFU, MUSE, and WMDP) and also enables analyses of forgetting behaviors across 450+ checkpoints we publicly release. Leveraging OpenUnlearning, we propose a novel meta-evaluation benchmark focused specifically on assessing the faithfulness and robustness of evaluation metrics themselves. We also benchmark diverse unlearning methods and provide a comparative analysis against an extensive evaluation suite. Overall, we establish a clear, community-driven pathway toward rigorous development in LLM unlearning research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12618
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
Dorna, Vineeth
Mekala, Anmol
Zhao, Wenlong
McCallum, Andrew
Lipton, Zachary C.
Kolter, J. Zico
Maini, Pratyush
Computation and Language
Robust unlearning is crucial for safely deploying large language models (LLMs) in environments where data privacy, model safety, and regulatory compliance must be ensured. Yet the task is inherently challenging, partly due to difficulties in reliably measuring whether unlearning has truly occurred. Moreover, fragmentation in current methodologies and inconsistent evaluation metrics hinder comparative analysis and reproducibility. To unify and accelerate research efforts, we introduce OpenUnlearning, a standardized and extensible framework designed explicitly for benchmarking both LLM unlearning methods and metrics. OpenUnlearning integrates 13 unlearning algorithms and 16 diverse evaluations across 3 leading benchmarks (TOFU, MUSE, and WMDP) and also enables analyses of forgetting behaviors across 450+ checkpoints we publicly release. Leveraging OpenUnlearning, we propose a novel meta-evaluation benchmark focused specifically on assessing the faithfulness and robustness of evaluation metrics themselves. We also benchmark diverse unlearning methods and provide a comparative analysis against an extensive evaluation suite. Overall, we establish a clear, community-driven pathway toward rigorous development in LLM unlearning research.
title OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
topic Computation and Language
url https://arxiv.org/abs/2506.12618