UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Lulu, Hu, Shengshan, Lu, Wei, Zhou, Ziqi, Song, Yufei, Cheng, Jianhong, Li, Minghui, Zhang, Yanjun, Zhang, Leo Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914286077476864
author Xue, Lulu
Hu, Shengshan
Lu, Wei
Zhou, Ziqi
Song, Yufei
Cheng, Jianhong
Li, Minghui
Zhang, Yanjun
Zhang, Leo Yu
author_facet Xue, Lulu
Hu, Shengshan
Lu, Wei
Zhou, Ziqi
Song, Yufei
Cheng, Jianhong
Li, Minghui
Zhang, Yanjun
Zhang, Leo Yu
contents Machine unlearning is an emerging technique that aims to remove the influence of specific data from trained models, thereby enhancing privacy protection. However, recent research has uncovered critical privacy vulnerabilities, showing that adversaries can exploit unlearning inversion to reconstruct data that was intended to be erased. Despite the severity of this threat, dedicated defenses remain lacking. To address this gap, we propose UnlearnShield, the first defense specifically tailored to counter unlearning inversion. UnlearnShield introduces directional perturbations in the cosine representation space and regulates them through a constraint module to jointly preserve model accuracy and forgetting efficacy, thereby reducing inversion risk while maintaining utility. Experiments demonstrate that it achieves a good trade-off among privacy protection, accuracy, and forgetting.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20325
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion
Xue, Lulu
Hu, Shengshan
Lu, Wei
Zhou, Ziqi
Song, Yufei
Cheng, Jianhong
Li, Minghui
Zhang, Yanjun
Zhang, Leo Yu
Cryptography and Security
Computer Vision and Pattern Recognition
Machine unlearning is an emerging technique that aims to remove the influence of specific data from trained models, thereby enhancing privacy protection. However, recent research has uncovered critical privacy vulnerabilities, showing that adversaries can exploit unlearning inversion to reconstruct data that was intended to be erased. Despite the severity of this threat, dedicated defenses remain lacking. To address this gap, we propose UnlearnShield, the first defense specifically tailored to counter unlearning inversion. UnlearnShield introduces directional perturbations in the cosine representation space and regulates them through a constraint module to jointly preserve model accuracy and forgetting efficacy, thereby reducing inversion risk while maintaining utility. Experiments demonstrate that it achieves a good trade-off among privacy protection, accuracy, and forgetting.
title UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion
topic Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.20325