Caption, Create, Continue: Continual Learning with Pre-trained Generative Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Solomon, Indu, Aung, Aye Phyu Phyu, Kumar, Uttam, Jayavelu, Senthilnath
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917076964212736
author Solomon, Indu
Aung, Aye Phyu Phyu
Kumar, Uttam
Jayavelu, Senthilnath
author_facet Solomon, Indu
Aung, Aye Phyu Phyu
Kumar, Uttam
Jayavelu, Senthilnath
contents Continual learning (CL) enables models to adapt to evolving data streams without catastrophic forgetting, a fundamental requirement for real-world AI systems. However, the current methods often depend on large replay buffers or heavily annotated datasets which are impractical due to storage, privacy, and cost constraints. We propose CLTS (Continual Learning via Text-Image Synergy), a novel class-incremental framework that mitigates forgetting without storing real task data. CLTS leverages pre-trained vision-language models, BLIP (Bootstrapping Language-Image Pre-training) for caption generation and stable diffusion for sample generation. Each task is handled by a dedicated Task Head, while a Task Router learns to assign inputs to the correct Task Head using the generated data. On three benchmark datasets, CLTS improves average task accuracy by up to 54% and achieves 63 times better memory efficiency compared to four recent continual learning baselines, demonstrating improved retention and adaptability. CLTS introduces a novel perspective by integrating generative text-image augmentation for scalable continual learning.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17806
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Caption, Create, Continue: Continual Learning with Pre-trained Generative Vision-Language Models
Solomon, Indu
Aung, Aye Phyu Phyu
Kumar, Uttam
Jayavelu, Senthilnath
Machine Learning
Continual learning (CL) enables models to adapt to evolving data streams without catastrophic forgetting, a fundamental requirement for real-world AI systems. However, the current methods often depend on large replay buffers or heavily annotated datasets which are impractical due to storage, privacy, and cost constraints. We propose CLTS (Continual Learning via Text-Image Synergy), a novel class-incremental framework that mitigates forgetting without storing real task data. CLTS leverages pre-trained vision-language models, BLIP (Bootstrapping Language-Image Pre-training) for caption generation and stable diffusion for sample generation. Each task is handled by a dedicated Task Head, while a Task Router learns to assign inputs to the correct Task Head using the generated data. On three benchmark datasets, CLTS improves average task accuracy by up to 54% and achieves 63 times better memory efficiency compared to four recent continual learning baselines, demonstrating improved retention and adaptability. CLTS introduces a novel perspective by integrating generative text-image augmentation for scalable continual learning.
title Caption, Create, Continue: Continual Learning with Pre-trained Generative Vision-Language Models
topic Machine Learning
url https://arxiv.org/abs/2409.17806