Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yu, Jiazuo, Zhuge, Yunzhi, Zhang, Lu, Hu, Ping, Wang, Dong, Lu, Huchuan, He, You
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929370856161280
author Yu, Jiazuo
Zhuge, Yunzhi
Zhang, Lu
Hu, Ping
Wang, Dong
Lu, Huchuan
He, You
author_facet Yu, Jiazuo
Zhuge, Yunzhi
Zhang, Lu
Hu, Ping
Wang, Dong
Lu, Huchuan
He, You
contents Continual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout lifelong learning and (ii) significant computational burdens associated with full-model tuning. In this work, we present a parameter-efficient continual learning framework to alleviate long-term forgetting in incremental learning with vision-language models. Our approach involves the dynamic expansion of a pre-trained CLIP model, through the integration of Mixture-of-Experts (MoE) adapters in response to new tasks. To preserve the zero-shot recognition capability of vision-language models, we further introduce a Distribution Discriminative Auto-Selector (DDAS) that automatically routes in-distribution and out-of-distribution inputs to the MoE Adapter and the original CLIP, respectively. Through extensive experiments across various settings, our proposed method consistently outperforms previous state-of-the-art approaches while concurrently reducing parameter training burdens by 60%. Our code locates at https://github.com/JiazuoYu/MoE-Adapters4CL
format Preprint
id arxiv_https___arxiv_org_abs_2403_11549
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
Yu, Jiazuo
Zhuge, Yunzhi
Zhang, Lu
Hu, Ping
Wang, Dong
Lu, Huchuan
He, You
Computer Vision and Pattern Recognition
Continual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout lifelong learning and (ii) significant computational burdens associated with full-model tuning. In this work, we present a parameter-efficient continual learning framework to alleviate long-term forgetting in incremental learning with vision-language models. Our approach involves the dynamic expansion of a pre-trained CLIP model, through the integration of Mixture-of-Experts (MoE) adapters in response to new tasks. To preserve the zero-shot recognition capability of vision-language models, we further introduce a Distribution Discriminative Auto-Selector (DDAS) that automatically routes in-distribution and out-of-distribution inputs to the MoE Adapter and the original CLIP, respectively. Through extensive experiments across various settings, our proposed method consistently outperforms previous state-of-the-art approaches while concurrently reducing parameter training burdens by 60%. Our code locates at https://github.com/JiazuoYu/MoE-Adapters4CL
title Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.11549