Towards Lifelong Few-Shot Customization of Text-to-Image Diffusion

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Song, Nan, Yang, Xiaofeng, Yang, Ze, Lin, Guosheng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910689979793408
author Song, Nan
Yang, Xiaofeng
Yang, Ze
Lin, Guosheng
author_facet Song, Nan
Yang, Xiaofeng
Yang, Ze
Lin, Guosheng
contents Lifelong few-shot customization for text-to-image diffusion aims to continually generalize existing models for new tasks with minimal data while preserving old knowledge. Current customization diffusion models excel in few-shot tasks but struggle with catastrophic forgetting problems in lifelong generations. In this study, we identify and categorize the catastrophic forgetting problems into two folds: relevant concepts forgetting and previous concepts forgetting. To address these challenges, we first devise a data-free knowledge distillation strategy to tackle relevant concepts forgetting. Unlike existing methods that rely on additional real data or offline replay of original concept data, our approach enables on-the-fly knowledge distillation to retain the previous concepts while learning new ones, without accessing any previous data. Second, we develop an In-Context Generation (ICGen) paradigm that allows the diffusion model to be conditioned upon the input vision context, which facilitates the few-shot generation and mitigates the issue of previous concepts forgetting. Extensive experiments show that the proposed Lifelong Few-Shot Diffusion (LFS-Diffusion) method can produce high-quality and accurate images while maintaining previously learned knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05544
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Lifelong Few-Shot Customization of Text-to-Image Diffusion
Song, Nan
Yang, Xiaofeng
Yang, Ze
Lin, Guosheng
Computer Vision and Pattern Recognition
Machine Learning
Lifelong few-shot customization for text-to-image diffusion aims to continually generalize existing models for new tasks with minimal data while preserving old knowledge. Current customization diffusion models excel in few-shot tasks but struggle with catastrophic forgetting problems in lifelong generations. In this study, we identify and categorize the catastrophic forgetting problems into two folds: relevant concepts forgetting and previous concepts forgetting. To address these challenges, we first devise a data-free knowledge distillation strategy to tackle relevant concepts forgetting. Unlike existing methods that rely on additional real data or offline replay of original concept data, our approach enables on-the-fly knowledge distillation to retain the previous concepts while learning new ones, without accessing any previous data. Second, we develop an In-Context Generation (ICGen) paradigm that allows the diffusion model to be conditioned upon the input vision context, which facilitates the few-shot generation and mitigates the issue of previous concepts forgetting. Extensive experiments show that the proposed Lifelong Few-Shot Diffusion (LFS-Diffusion) method can produce high-quality and accurate images while maintaining previously learned knowledge.
title Towards Lifelong Few-Shot Customization of Text-to-Image Diffusion
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.05544