Machine Unlearning under Retain-Forget Entanglement

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cheng, Jingpu, Liu, Ping, Li, Qianxiao, Zhang, Chi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911548386050048
author Cheng, Jingpu
Liu, Ping
Li, Qianxiao
Zhang, Chi
author_facet Cheng, Jingpu
Liu, Ping
Li, Qianxiao
Zhang, Chi
contents Forgetting a subset in machine unlearning is rarely an isolated task. Often, retained samples that are closely related to the forget set can be unintentionally affected, particularly when they share correlated features from pretraining or exhibit strong semantic similarities. To address this challenge, we propose a novel two-phase optimization framework specifically designed to handle such retai-forget entanglements. In the first phase, an augmented Lagrangian method increases the loss on the forget set while preserving accuracy on less-related retained samples. The second phase applies a gradient projection step, regularized by the Wasserstein-2 distance, to mitigate performance degradation on semantically related retained samples without compromising the unlearning objective. We validate our approach through comprehensive experiments on multiple unlearning tasks, standard benchmark datasets, and diverse neural architectures, demonstrating that it achieves effective and reliable unlearning while outperforming existing baselines in both accuracy retention and removal fidelity.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26569
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Machine Unlearning under Retain-Forget Entanglement
Cheng, Jingpu
Liu, Ping
Li, Qianxiao
Zhang, Chi
Machine Learning
Forgetting a subset in machine unlearning is rarely an isolated task. Often, retained samples that are closely related to the forget set can be unintentionally affected, particularly when they share correlated features from pretraining or exhibit strong semantic similarities. To address this challenge, we propose a novel two-phase optimization framework specifically designed to handle such retai-forget entanglements. In the first phase, an augmented Lagrangian method increases the loss on the forget set while preserving accuracy on less-related retained samples. The second phase applies a gradient projection step, regularized by the Wasserstein-2 distance, to mitigate performance degradation on semantically related retained samples without compromising the unlearning objective. We validate our approach through comprehensive experiments on multiple unlearning tasks, standard benchmark datasets, and diverse neural architectures, demonstrating that it achieves effective and reliable unlearning while outperforming existing baselines in both accuracy retention and removal fidelity.
title Machine Unlearning under Retain-Forget Entanglement
topic Machine Learning
url https://arxiv.org/abs/2603.26569