WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maheri, Mohammad M, Cadet, Xavier, Chin, Peter, Haddadi, Hamed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910038570827776
author Maheri, Mohammad M
Cadet, Xavier
Chin, Peter
Haddadi, Hamed
author_facet Maheri, Mohammad M
Cadet, Xavier
Chin, Peter
Haddadi, Hamed
contents Approximate machine unlearning aims to efficiently remove the influence of specific data points from a trained model, offering a practical alternative to full retraining. However, it introduces privacy risks: an adversary with access to pre- and post-unlearning models can exploit their differences for membership inference or data reconstruction. We show these vulnerabilities arise from two factors: large gradient norms of forget-set samples and the close proximity of unlearned parameters to the original model. To demonstrate their severity, we propose unlearning-specific membership inference and reconstruction attacks, showing that several state-of-the-art methods (e.g., NGP, SCRUB) remain vulnerable. To mitigate this leakage, we introduce WARP, a plug-and-play teleportation defense that leverages neural network symmetries to reduce forget-set gradient energy and increase parameter dispersion while preserving predictions. This reparameterization obfuscates the signal of forgotten data, making it harder for attackers to distinguish forgotten samples from non-members or recover them via reconstruction. Across six unlearning algorithms, our approach achieves consistent privacy gains, reducing adversarial advantage (AUC) by up to 64% in black-box and 92% in white-box settings, while maintaining accuracy on retained data. These results highlight teleportation as a general tool for reducing attack success in approximate unlearning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00272
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols
Maheri, Mohammad M
Cadet, Xavier
Chin, Peter
Haddadi, Hamed
Machine Learning
Artificial Intelligence
Cryptography and Security
Approximate machine unlearning aims to efficiently remove the influence of specific data points from a trained model, offering a practical alternative to full retraining. However, it introduces privacy risks: an adversary with access to pre- and post-unlearning models can exploit their differences for membership inference or data reconstruction. We show these vulnerabilities arise from two factors: large gradient norms of forget-set samples and the close proximity of unlearned parameters to the original model. To demonstrate their severity, we propose unlearning-specific membership inference and reconstruction attacks, showing that several state-of-the-art methods (e.g., NGP, SCRUB) remain vulnerable. To mitigate this leakage, we introduce WARP, a plug-and-play teleportation defense that leverages neural network symmetries to reduce forget-set gradient energy and increase parameter dispersion while preserving predictions. This reparameterization obfuscates the signal of forgotten data, making it harder for attackers to distinguish forgotten samples from non-members or recover them via reconstruction. Across six unlearning algorithms, our approach achieves consistent privacy gains, reducing adversarial advantage (AUC) by up to 64% in black-box and 92% in white-box settings, while maintaining accuracy on retained data. These results highlight teleportation as a general tool for reducing attack success in approximate unlearning.
title WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2512.00272