PRUNE: A Patching Based Repair Framework for Certifiable Unlearning of Neural Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Xuran, Wang, Jingyi, Yuan, Xiaohan, Zhang, Peixin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918165942894592
author Li, Xuran
Wang, Jingyi
Yuan, Xiaohan
Zhang, Peixin
author_facet Li, Xuran
Wang, Jingyi
Yuan, Xiaohan
Zhang, Peixin
contents It is often desirable to remove (a.k.a. unlearn) a specific part of the training data from a trained neural network model. A typical application scenario is to protect the data holder's right to be forgotten, which has been promoted by many recent regulation rules. Existing unlearning methods involve training alternative models with remaining data, which may be costly and challenging to verify from the data holder or a thirdparty auditor's perspective. In this work, we provide a new angle and propose a novel unlearning approach by imposing carefully crafted "patch" on the original neural network to achieve targeted "forgetting" of the requested data to delete. Specifically, inspired by the research line of neural network repair, we propose to strategically seek a lightweight minimum "patch" for unlearning a given data point with certifiable guarantee. Furthermore, to unlearn a considerable amount of data points (or an entire class), we propose to iteratively select a small subset of representative data points to unlearn, which achieves the effect of unlearning the whole set. Extensive experiments on multiple categorical datasets demonstrates our approach's effectiveness, achieving measurable unlearning while preserving the model's performance and being competitive in efficiency and memory consumption compared to various baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2505_06520
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PRUNE: A Patching Based Repair Framework for Certifiable Unlearning of Neural Networks
Li, Xuran
Wang, Jingyi
Yuan, Xiaohan
Zhang, Peixin
Machine Learning
Artificial Intelligence
Cryptography and Security
It is often desirable to remove (a.k.a. unlearn) a specific part of the training data from a trained neural network model. A typical application scenario is to protect the data holder's right to be forgotten, which has been promoted by many recent regulation rules. Existing unlearning methods involve training alternative models with remaining data, which may be costly and challenging to verify from the data holder or a thirdparty auditor's perspective. In this work, we provide a new angle and propose a novel unlearning approach by imposing carefully crafted "patch" on the original neural network to achieve targeted "forgetting" of the requested data to delete. Specifically, inspired by the research line of neural network repair, we propose to strategically seek a lightweight minimum "patch" for unlearning a given data point with certifiable guarantee. Furthermore, to unlearn a considerable amount of data points (or an entire class), we propose to iteratively select a small subset of representative data points to unlearn, which achieves the effect of unlearning the whole set. Extensive experiments on multiple categorical datasets demonstrates our approach's effectiveness, achieving measurable unlearning while preserving the model's performance and being competitive in efficiency and memory consumption compared to various baseline methods.
title PRUNE: A Patching Based Repair Framework for Certifiable Unlearning of Neural Networks
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2505.06520