CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Frascaroli, Emanuele, Panariello, Aniello, Buzzega, Pietro, Bonicelli, Lorenzo, Porrello, Angelo, Calderara, Simone
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917819037253632
author Frascaroli, Emanuele
Panariello, Aniello
Buzzega, Pietro
Bonicelli, Lorenzo
Porrello, Angelo
Calderara, Simone
author_facet Frascaroli, Emanuele
Panariello, Aniello
Buzzega, Pietro
Bonicelli, Lorenzo
Porrello, Angelo
Calderara, Simone
contents With the emergence of Transformers and Vision-Language Models (VLMs) such as CLIP, fine-tuning large pre-trained models has recently become a prevalent strategy in Continual Learning. This has led to the development of numerous prompting strategies to adapt transformer-based models without incurring catastrophic forgetting. However, these strategies often compromise the original zero-shot capabilities of the pre-trained CLIP model and struggle to adapt to domains that significantly deviate from the pre-training data. In this work, we propose Continual Generative training for Incremental prompt-Learning, a simple and novel approach to mitigate forgetting while adapting CLIP. Briefly, we employ Variational Autoencoders (VAEs) to learn class-conditioned distributions within the embedding space of the visual encoder. We then exploit these distributions to sample new synthetic visual embeddings and train the corresponding class-specific textual prompts during subsequent tasks. Through extensive experiments on different domains, we show that such a generative replay approach can adapt to new tasks while improving zero-shot capabilities, evaluated using a novel metric tailored for CL scenarios. Notably, further analysis reveals that our approach can bridge the gap with joint prompt tuning. The codebase is available at https://github.com/aimagelab/mammoth.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15793
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning
Frascaroli, Emanuele
Panariello, Aniello
Buzzega, Pietro
Bonicelli, Lorenzo
Porrello, Angelo
Calderara, Simone
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
With the emergence of Transformers and Vision-Language Models (VLMs) such as CLIP, fine-tuning large pre-trained models has recently become a prevalent strategy in Continual Learning. This has led to the development of numerous prompting strategies to adapt transformer-based models without incurring catastrophic forgetting. However, these strategies often compromise the original zero-shot capabilities of the pre-trained CLIP model and struggle to adapt to domains that significantly deviate from the pre-training data. In this work, we propose Continual Generative training for Incremental prompt-Learning, a simple and novel approach to mitigate forgetting while adapting CLIP. Briefly, we employ Variational Autoencoders (VAEs) to learn class-conditioned distributions within the embedding space of the visual encoder. We then exploit these distributions to sample new synthetic visual embeddings and train the corresponding class-specific textual prompts during subsequent tasks. Through extensive experiments on different domains, we show that such a generative replay approach can adapt to new tasks while improving zero-shot capabilities, evaluated using a novel metric tailored for CL scenarios. Notably, further analysis reveals that our approach can bridge the gap with joint prompt tuning. The codebase is available at https://github.com/aimagelab/mammoth.
title CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.15793