Class-Incremental Learning with CLIP: Adaptive Representation Adjustment and Parameter Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Linlan, Cao, Xusheng, Lu, Haori, Liu, Xialei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910535099875328
author Huang, Linlan
Cao, Xusheng
Lu, Haori
Liu, Xialei
author_facet Huang, Linlan
Cao, Xusheng
Lu, Haori
Liu, Xialei
contents Class-incremental learning is a challenging problem, where the goal is to train a model that can classify data from an increasing number of classes over time. With the advancement of vision-language pre-trained models such as CLIP, they demonstrate good generalization ability that allows them to excel in class-incremental learning with completely frozen parameters. However, further adaptation to downstream tasks by simply fine-tuning the model leads to severe forgetting. Most existing works with pre-trained models assume that the forgetting of old classes is uniform when the model acquires new knowledge. In this paper, we propose a method named Adaptive Representation Adjustment and Parameter Fusion (RAPF). During training for new data, we measure the influence of new classes on old ones and adjust the representations, using textual features. After training, we employ a decomposed parameter fusion to further mitigate forgetting during adapter module fine-tuning. Experiments on several conventional benchmarks show that our method achieves state-of-the-art results. Our code is available at \url{https://github.com/linlany/RAPF}.
format Preprint
id arxiv_https___arxiv_org_abs_2407_14143
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Class-Incremental Learning with CLIP: Adaptive Representation Adjustment and Parameter Fusion
Huang, Linlan
Cao, Xusheng
Lu, Haori
Liu, Xialei
Computer Vision and Pattern Recognition
Machine Learning
Class-incremental learning is a challenging problem, where the goal is to train a model that can classify data from an increasing number of classes over time. With the advancement of vision-language pre-trained models such as CLIP, they demonstrate good generalization ability that allows them to excel in class-incremental learning with completely frozen parameters. However, further adaptation to downstream tasks by simply fine-tuning the model leads to severe forgetting. Most existing works with pre-trained models assume that the forgetting of old classes is uniform when the model acquires new knowledge. In this paper, we propose a method named Adaptive Representation Adjustment and Parameter Fusion (RAPF). During training for new data, we measure the influence of new classes on old ones and adjust the representations, using textual features. After training, we employ a decomposed parameter fusion to further mitigate forgetting during adapter module fine-tuning. Experiments on several conventional benchmarks show that our method achieves state-of-the-art results. Our code is available at \url{https://github.com/linlany/RAPF}.
title Class-Incremental Learning with CLIP: Adaptive Representation Adjustment and Parameter Fusion
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2407.14143