Understanding Fine-tuning in Approximate Unlearning: A Theoretical Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Meng, Sharma, Rohan, Chen, Changyou, Xu, Jinhui, Ji, Kaiyi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915631404679168
author Ding, Meng
Sharma, Rohan
Chen, Changyou
Xu, Jinhui
Ji, Kaiyi
author_facet Ding, Meng
Sharma, Rohan
Chen, Changyou
Xu, Jinhui
Ji, Kaiyi
contents Machine Unlearning has emerged as a significant area of research, focusing on `removing' specific subsets of data from a trained model. Fine-tuning (FT) methods have become one of the fundamental approaches for approximating unlearning, as they effectively retain model performance. However, it is consistently observed that naive FT methods struggle to forget the targeted data. In this paper, we present the first theoretical analysis of FT methods for machine unlearning within a linear regression framework, providing a deeper exploration of this phenomenon. Our analysis reveals that while FT models can achieve zero remaining loss, they fail to forget the forgetting data, as the pretrained model retains its influence and the fine-tuning process does not adequately mitigate it. To address this, we propose a novel Retention-Based Masking (RBM) strategy that constructs a weight saliency map based on the remaining dataset, unlike existing methods that focus on the forgetting dataset. Our theoretical analysis demonstrates that RBM not only significantly improves unlearning accuracy (UA) but also ensures higher retaining accuracy (RA) by preserving overlapping features shared between the forgetting and remaining datasets. Experiments on synthetic and real-world datasets validate our theoretical insights, showing that RBM outperforms existing masking approaches in balancing UA, RA, and disparity metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03833
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding Fine-tuning in Approximate Unlearning: A Theoretical Perspective
Ding, Meng
Sharma, Rohan
Chen, Changyou
Xu, Jinhui
Ji, Kaiyi
Machine Learning
Machine Unlearning has emerged as a significant area of research, focusing on `removing' specific subsets of data from a trained model. Fine-tuning (FT) methods have become one of the fundamental approaches for approximating unlearning, as they effectively retain model performance. However, it is consistently observed that naive FT methods struggle to forget the targeted data. In this paper, we present the first theoretical analysis of FT methods for machine unlearning within a linear regression framework, providing a deeper exploration of this phenomenon. Our analysis reveals that while FT models can achieve zero remaining loss, they fail to forget the forgetting data, as the pretrained model retains its influence and the fine-tuning process does not adequately mitigate it. To address this, we propose a novel Retention-Based Masking (RBM) strategy that constructs a weight saliency map based on the remaining dataset, unlike existing methods that focus on the forgetting dataset. Our theoretical analysis demonstrates that RBM not only significantly improves unlearning accuracy (UA) but also ensures higher retaining accuracy (RA) by preserving overlapping features shared between the forgetting and remaining datasets. Experiments on synthetic and real-world datasets validate our theoretical insights, showing that RBM outperforms existing masking approaches in balancing UA, RA, and disparity metrics.
title Understanding Fine-tuning in Approximate Unlearning: A Theoretical Perspective
topic Machine Learning
url https://arxiv.org/abs/2410.03833