CLIP model is an Efficient Online Lifelong Learner

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Leyuan, Xiang, Liuyu, Wei, Yujie, Wang, Yunlong, He, Zhaofeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929357369376768
author Wang, Leyuan
Xiang, Liuyu
Wei, Yujie
Wang, Yunlong
He, Zhaofeng
author_facet Wang, Leyuan
Xiang, Liuyu
Wei, Yujie
Wang, Yunlong
He, Zhaofeng
contents Online Lifelong Learning (OLL) addresses the challenge of learning from continuous and non-stationary data streams. Existing online lifelong learning methods based on image classification models often require preset conditions such as the total number of classes or maximum memory capacity, which hinders the realization of real never-ending learning and renders them impractical for real-world scenarios. In this work, we propose that vision-language models, such as Contrastive Language-Image Pretraining (CLIP), are more suitable candidates for online lifelong learning. We discover that maintaining symmetry between image and text is crucial during Parameter-Efficient Tuning (PET) for CLIP model in online lifelong learning. To this end, we introduce the Symmetric Image-Text (SIT) tuning strategy. We conduct extensive experiments on multiple lifelong learning benchmark datasets and elucidate the effectiveness of SIT through gradient analysis. Additionally, we assess the impact of lifelong learning on generalizability of CLIP and found that tuning the image encoder is beneficial for lifelong learning, while tuning the text encoder aids in zero-shot learning.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15155
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLIP model is an Efficient Online Lifelong Learner
Wang, Leyuan
Xiang, Liuyu
Wei, Yujie
Wang, Yunlong
He, Zhaofeng
Computer Vision and Pattern Recognition
Online Lifelong Learning (OLL) addresses the challenge of learning from continuous and non-stationary data streams. Existing online lifelong learning methods based on image classification models often require preset conditions such as the total number of classes or maximum memory capacity, which hinders the realization of real never-ending learning and renders them impractical for real-world scenarios. In this work, we propose that vision-language models, such as Contrastive Language-Image Pretraining (CLIP), are more suitable candidates for online lifelong learning. We discover that maintaining symmetry between image and text is crucial during Parameter-Efficient Tuning (PET) for CLIP model in online lifelong learning. To this end, we introduce the Symmetric Image-Text (SIT) tuning strategy. We conduct extensive experiments on multiple lifelong learning benchmark datasets and elucidate the effectiveness of SIT through gradient analysis. Additionally, we assess the impact of lifelong learning on generalizability of CLIP and found that tuning the image encoder is beneficial for lifelong learning, while tuning the text encoder aids in zero-shot learning.
title CLIP model is an Efficient Online Lifelong Learner
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15155