LVP-CLIP:Revisiting CLIP for Continual Learning with Label Vector Pool

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ma, Yue, Ren, Huantao, Wang, Boyu, Jin, Jingang, Velipasalar, Senem, Qiu, Qinru
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909419439128576
author Ma, Yue
Ren, Huantao
Wang, Boyu
Jin, Jingang
Velipasalar, Senem
Qiu, Qinru
author_facet Ma, Yue
Ren, Huantao
Wang, Boyu
Jin, Jingang
Velipasalar, Senem
Qiu, Qinru
contents Continual learning aims to update a model so that it can sequentially learn new tasks without forgetting previously acquired knowledge. Recent continual learning approaches often leverage the vision-language model CLIP for its high-dimensional feature space and cross-modality feature matching. Traditional CLIP-based classification methods identify the most similar text label for a test image by comparing their embeddings. However, these methods are sensitive to the quality of text phrases and less effective for classes lacking meaningful text labels. In this work, we rethink CLIP-based continual learning and introduce the concept of Label Vector Pool (LVP). LVP replaces text labels with training images as similarity references, eliminating the need for ideal text descriptions. We present three variations of LVP and evaluate their performance on class and domain incremental learning tasks. Leveraging CLIP's high dimensional feature space, LVP learning algorithms are task-order invariant. The new knowledge does not modify the old knowledge, hence, there is minimum forgetting. Different tasks can be learned independently and in parallel with low computational and memory demands. Experimental results show that proposed LVP-based methods outperform the current state-of-the-art baseline by a significant margin of 40.7%.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05840
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LVP-CLIP:Revisiting CLIP for Continual Learning with Label Vector Pool
Ma, Yue
Ren, Huantao
Wang, Boyu
Jin, Jingang
Velipasalar, Senem
Qiu, Qinru
Computer Vision and Pattern Recognition
68T45
I.2.10; I.4; I.5
Continual learning aims to update a model so that it can sequentially learn new tasks without forgetting previously acquired knowledge. Recent continual learning approaches often leverage the vision-language model CLIP for its high-dimensional feature space and cross-modality feature matching. Traditional CLIP-based classification methods identify the most similar text label for a test image by comparing their embeddings. However, these methods are sensitive to the quality of text phrases and less effective for classes lacking meaningful text labels. In this work, we rethink CLIP-based continual learning and introduce the concept of Label Vector Pool (LVP). LVP replaces text labels with training images as similarity references, eliminating the need for ideal text descriptions. We present three variations of LVP and evaluate their performance on class and domain incremental learning tasks. Leveraging CLIP's high dimensional feature space, LVP learning algorithms are task-order invariant. The new knowledge does not modify the old knowledge, hence, there is minimum forgetting. Different tasks can be learned independently and in parallel with low computational and memory demands. Experimental results show that proposed LVP-based methods outperform the current state-of-the-art baseline by a significant margin of 40.7%.
title LVP-CLIP:Revisiting CLIP for Continual Learning with Label Vector Pool
topic Computer Vision and Pattern Recognition
68T45
I.2.10; I.4; I.5
url https://arxiv.org/abs/2412.05840