Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cha, Sungmin, Cho, Sungjun, Hwang, Dasol, Lee, Honglak, Moon, Taesup, Lee, Moontae
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929209173082112
author Cha, Sungmin
Cho, Sungjun
Hwang, Dasol
Lee, Honglak
Moon, Taesup
Lee, Moontae
author_facet Cha, Sungmin
Cho, Sungjun
Hwang, Dasol
Lee, Honglak
Moon, Taesup
Lee, Moontae
contents Since the recent advent of regulations for data protection (e.g., the General Data Protection Regulation), there has been increasing demand in deleting information learned from sensitive data in pre-trained models without retraining from scratch. The inherent vulnerability of neural networks towards adversarial attacks and unfairness also calls for a robust method to remove or correct information in an instance-wise fashion, while retaining the predictive performance across remaining data. To this end, we consider instance-wise unlearning, of which the goal is to delete information on a set of instances from a pre-trained model, by either misclassifying each instance away from its original prediction or relabeling the instance to a different label. We also propose two methods that reduce forgetting on the remaining data: 1) utilizing adversarial examples to overcome forgetting at the representation-level and 2) leveraging weight importance metrics to pinpoint network parameters guilty of propagating unwanted information. Both methods only require the pre-trained model and data instances to forget, allowing painless application to real-life settings where the entire training set is unavailable. Through extensive experimentation on various image classification benchmarks, we show that our approach effectively preserves knowledge of remaining data while unlearning given instances in both single-task and continual unlearning scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2301_11578
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
Cha, Sungmin
Cho, Sungjun
Hwang, Dasol
Lee, Honglak
Moon, Taesup
Lee, Moontae
Machine Learning
Since the recent advent of regulations for data protection (e.g., the General Data Protection Regulation), there has been increasing demand in deleting information learned from sensitive data in pre-trained models without retraining from scratch. The inherent vulnerability of neural networks towards adversarial attacks and unfairness also calls for a robust method to remove or correct information in an instance-wise fashion, while retaining the predictive performance across remaining data. To this end, we consider instance-wise unlearning, of which the goal is to delete information on a set of instances from a pre-trained model, by either misclassifying each instance away from its original prediction or relabeling the instance to a different label. We also propose two methods that reduce forgetting on the remaining data: 1) utilizing adversarial examples to overcome forgetting at the representation-level and 2) leveraging weight importance metrics to pinpoint network parameters guilty of propagating unwanted information. Both methods only require the pre-trained model and data instances to forget, allowing painless application to real-life settings where the entire training set is unavailable. Through extensive experimentation on various image classification benchmarks, we show that our approach effectively preserves knowledge of remaining data while unlearning given instances in both single-task and continual unlearning scenarios.
title Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
topic Machine Learning
url https://arxiv.org/abs/2301.11578