Machine Unlearning on Pre-trained Models by Residual Feature Alignment Using LoRA

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qin, Laiqiao, Zhu, Tianqing, Wang, Linlin, Zhou, Wanlei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915996399304704
author Qin, Laiqiao
Zhu, Tianqing
Wang, Linlin
Zhou, Wanlei
author_facet Qin, Laiqiao
Zhu, Tianqing
Wang, Linlin
Zhou, Wanlei
contents Machine unlearning is an emerging technology that removes a subset of the training data from a trained model without significantly affecting the model performance on the remaining data. This topic is becoming increasingly important in protecting user privacy and eliminating harmful or outdated data. The key challenge lies in effectively and efficiently unlearning specific information without compromising the model's utility on the retained data. For pre-trained models, fine-tuning is an important way to achieve the unlearning target. Previous work typically fine-tuned the entire model's parameters, which incurred significant computational costs. In addition, the fine-tuning process may cause shifts in the intermediate layer features, affecting the model's overall utility. In this work, we propose a novel and efficient machine unlearning method for pre-trained models. We term the method Residual Feature Alignment Unlearning. Specifically, we leverage LoRA (Low-Rank Adaptation) to decompose the model's intermediate features into pre-trained features and residual features. By adjusting the residual features, we align the unlearned model with the pre-trained model at the intermediate feature level to achieve both unlearning and remaining targets. The method aims to learn zero residuals on the retained set and shifted residuals on the unlearning set. Extensive experiments on numerous datasets validate the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08443
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Machine Unlearning on Pre-trained Models by Residual Feature Alignment Using LoRA
Qin, Laiqiao
Zhu, Tianqing
Wang, Linlin
Zhou, Wanlei
Machine Learning
Computer Vision and Pattern Recognition
Machine unlearning is an emerging technology that removes a subset of the training data from a trained model without significantly affecting the model performance on the remaining data. This topic is becoming increasingly important in protecting user privacy and eliminating harmful or outdated data. The key challenge lies in effectively and efficiently unlearning specific information without compromising the model's utility on the retained data. For pre-trained models, fine-tuning is an important way to achieve the unlearning target. Previous work typically fine-tuned the entire model's parameters, which incurred significant computational costs. In addition, the fine-tuning process may cause shifts in the intermediate layer features, affecting the model's overall utility. In this work, we propose a novel and efficient machine unlearning method for pre-trained models. We term the method Residual Feature Alignment Unlearning. Specifically, we leverage LoRA (Low-Rank Adaptation) to decompose the model's intermediate features into pre-trained features and residual features. By adjusting the residual features, we align the unlearned model with the pre-trained model at the intermediate feature level to achieve both unlearning and remaining targets. The method aims to learn zero residuals on the retained set and shifted residuals on the unlearning set. Extensive experiments on numerous datasets validate the effectiveness of our approach.
title Machine Unlearning on Pre-trained Models by Residual Feature Alignment Using LoRA
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.08443