CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Tianyu, Dai, Lisen, Wang, Xiangqi, Cheng, Minhao, Tian, Yapeng, Zhang, Xiangliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910990774304768
author Yang, Tianyu
Dai, Lisen
Wang, Xiangqi
Cheng, Minhao
Tian, Yapeng
Zhang, Xiangliang
author_facet Yang, Tianyu
Dai, Lisen
Wang, Xiangqi
Cheng, Minhao
Tian, Yapeng
Zhang, Xiangliang
contents Machine unlearning (MU) has gained significant attention as a means to remove specific data from trained models without requiring a full retraining process. While progress has been made in unimodal domains like text and image classification, unlearning in multimodal models remains relatively underexplored. In this work, we address the unique challenges of unlearning in CLIP, a prominent multimodal model that aligns visual and textual representations. We introduce CLIPErase, a novel approach that disentangles and selectively forgets both visual and textual associations, ensuring that unlearning does not compromise model performance. CLIPErase consists of three key modules: a Forgetting Module that disrupts the associations in the forget set, a Retention Module that preserves performance on the retain set, and a Consistency Module that maintains consistency with the original model. Extensive experiments on the CIFAR-100 and Flickr30K datasets across four CLIP downstream tasks demonstrate that CLIPErase effectively forgets designated associations in zero-shot tasks for multimodal samples, while preserving the model's performance on the retain set after unlearning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23330
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
Yang, Tianyu
Dai, Lisen
Wang, Xiangqi
Cheng, Minhao
Tian, Yapeng
Zhang, Xiangliang
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Machine unlearning (MU) has gained significant attention as a means to remove specific data from trained models without requiring a full retraining process. While progress has been made in unimodal domains like text and image classification, unlearning in multimodal models remains relatively underexplored. In this work, we address the unique challenges of unlearning in CLIP, a prominent multimodal model that aligns visual and textual representations. We introduce CLIPErase, a novel approach that disentangles and selectively forgets both visual and textual associations, ensuring that unlearning does not compromise model performance. CLIPErase consists of three key modules: a Forgetting Module that disrupts the associations in the forget set, a Retention Module that preserves performance on the retain set, and a Consistency Module that maintains consistency with the original model. Extensive experiments on the CIFAR-100 and Flickr30K datasets across four CLIP downstream tasks demonstrate that CLIPErase effectively forgets designated associations in zero-shot tasks for multimodal samples, while preserving the model's performance on the retain set after unlearning.
title CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.23330