Feature Unlearning for Pre-trained GANs and VAEs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moon, Saemi, Cho, Seunghyuk, Kim, Dongwoo
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917624422596608
author Moon, Saemi
Cho, Seunghyuk
Kim, Dongwoo
author_facet Moon, Saemi
Cho, Seunghyuk
Kim, Dongwoo
contents We tackle the problem of feature unlearning from a pre-trained image generative model: GANs and VAEs. Unlike a common unlearning task where an unlearning target is a subset of the training set, we aim to unlearn a specific feature, such as hairstyle from facial images, from the pre-trained generative models. As the target feature is only presented in a local region of an image, unlearning the entire image from the pre-trained model may result in losing other details in the remaining region of the image. To specify which features to unlearn, we collect randomly generated images that contain the target features. We then identify a latent representation corresponding to the target feature and then use the representation to fine-tune the pre-trained model. Through experiments on MNIST, CelebA, and FFHQ datasets, we show that target features are successfully removed while keeping the fidelity of the original models. Further experiments with an adversarial attack show that the unlearned model is more robust under the presence of malicious parties.
format Preprint
id arxiv_https___arxiv_org_abs_2303_05699
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Feature Unlearning for Pre-trained GANs and VAEs
Moon, Saemi
Cho, Seunghyuk
Kim, Dongwoo
Computer Vision and Pattern Recognition
Machine Learning
We tackle the problem of feature unlearning from a pre-trained image generative model: GANs and VAEs. Unlike a common unlearning task where an unlearning target is a subset of the training set, we aim to unlearn a specific feature, such as hairstyle from facial images, from the pre-trained generative models. As the target feature is only presented in a local region of an image, unlearning the entire image from the pre-trained model may result in losing other details in the remaining region of the image. To specify which features to unlearn, we collect randomly generated images that contain the target features. We then identify a latent representation corresponding to the target feature and then use the representation to fine-tune the pre-trained model. Through experiments on MNIST, CelebA, and FFHQ datasets, we show that target features are successfully removed while keeping the fidelity of the original models. Further experiments with an adversarial attack show that the unlearned model is more robust under the presence of malicious parties.
title Feature Unlearning for Pre-trained GANs and VAEs
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2303.05699