Provable Contrastive Continual Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wen, Yichen, Tan, Zhiquan, Zheng, Kaipeng, Xie, Chuanlong, Huang, Weiran
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916264948006912
author Wen, Yichen
Tan, Zhiquan
Zheng, Kaipeng
Xie, Chuanlong
Huang, Weiran
author_facet Wen, Yichen
Tan, Zhiquan
Zheng, Kaipeng
Xie, Chuanlong
Huang, Weiran
contents Continual learning requires learning incremental tasks with dynamic data distributions. So far, it has been observed that employing a combination of contrastive loss and distillation loss for training in continual learning yields strong performance. To the best of our knowledge, however, this contrastive continual learning framework lacks convincing theoretical explanations. In this work, we fill this gap by establishing theoretical performance guarantees, which reveal how the performance of the model is bounded by training losses of previous tasks in the contrastive continual learning framework. Our theoretical explanations further support the idea that pre-training can benefit continual learning. Inspired by our theoretical analysis of these guarantees, we propose a novel contrastive continual learning algorithm called CILA, which uses adaptive distillation coefficients for different tasks. These distillation coefficients are easily computed by the ratio between average distillation losses and average contrastive losses from previous tasks. Our method shows great improvement on standard benchmarks and achieves new state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18756
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Provable Contrastive Continual Learning
Wen, Yichen
Tan, Zhiquan
Zheng, Kaipeng
Xie, Chuanlong
Huang, Weiran
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Applications
Continual learning requires learning incremental tasks with dynamic data distributions. So far, it has been observed that employing a combination of contrastive loss and distillation loss for training in continual learning yields strong performance. To the best of our knowledge, however, this contrastive continual learning framework lacks convincing theoretical explanations. In this work, we fill this gap by establishing theoretical performance guarantees, which reveal how the performance of the model is bounded by training losses of previous tasks in the contrastive continual learning framework. Our theoretical explanations further support the idea that pre-training can benefit continual learning. Inspired by our theoretical analysis of these guarantees, we propose a novel contrastive continual learning algorithm called CILA, which uses adaptive distillation coefficients for different tasks. These distillation coefficients are easily computed by the ratio between average distillation losses and average contrastive losses from previous tasks. Our method shows great improvement on standard benchmarks and achieves new state-of-the-art performance.
title Provable Contrastive Continual Learning
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Applications
url https://arxiv.org/abs/2405.18756