Machine Unlearning under Retain-Forget Entanglement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Jingpu, Liu, Ping, Li, Qianxiao, Zhang, Chi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911548386050048
author Cheng, Jingpu
Liu, Ping
Li, Qianxiao
Zhang, Chi
author_facet Cheng, Jingpu
Liu, Ping
Li, Qianxiao
Zhang, Chi
contents Forgetting a subset in machine unlearning is rarely an isolated task. Often, retained samples that are closely related to the forget set can be unintentionally affected, particularly when they share correlated features from pretraining or exhibit strong semantic similarities. To address this challenge, we propose a novel two-phase optimization framework specifically designed to handle such retai-forget entanglements. In the first phase, an augmented Lagrangian method increases the loss on the forget set while preserving accuracy on less-related retained samples. The second phase applies a gradient projection step, regularized by the Wasserstein-2 distance, to mitigate performance degradation on semantically related retained samples without compromising the unlearning objective. We validate our approach through comprehensive experiments on multiple unlearning tasks, standard benchmark datasets, and diverse neural architectures, demonstrating that it achieves effective and reliable unlearning while outperforming existing baselines in both accuracy retention and removal fidelity.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26569
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Machine Unlearning under Retain-Forget Entanglement
Cheng, Jingpu
Liu, Ping
Li, Qianxiao
Zhang, Chi
Machine Learning
Forgetting a subset in machine unlearning is rarely an isolated task. Often, retained samples that are closely related to the forget set can be unintentionally affected, particularly when they share correlated features from pretraining or exhibit strong semantic similarities. To address this challenge, we propose a novel two-phase optimization framework specifically designed to handle such retai-forget entanglements. In the first phase, an augmented Lagrangian method increases the loss on the forget set while preserving accuracy on less-related retained samples. The second phase applies a gradient projection step, regularized by the Wasserstein-2 distance, to mitigate performance degradation on semantically related retained samples without compromising the unlearning objective. We validate our approach through comprehensive experiments on multiple unlearning tasks, standard benchmark datasets, and diverse neural architectures, demonstrating that it achieves effective and reliable unlearning while outperforming existing baselines in both accuracy retention and removal fidelity.
title Machine Unlearning under Retain-Forget Entanglement
topic Machine Learning
url https://arxiv.org/abs/2603.26569