VidCLearn: A Continual Learning Approach for Text-to-Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zanchetta, Luca, Papa, Lorenzo, Maiano, Luca, Amerini, Irene
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915504824778752
author Zanchetta, Luca
Papa, Lorenzo
Maiano, Luca
Amerini, Irene
author_facet Zanchetta, Luca
Papa, Lorenzo
Maiano, Luca
Amerini, Irene
contents Text-to-video generation is an emerging field in generative AI, enabling the creation of realistic, semantically accurate videos from text prompts. While current models achieve impressive visual quality and alignment with input text, they typically rely on static knowledge, making it difficult to incorporate new data without retraining from scratch. To address this limitation, we propose VidCLearn, a continual learning framework for diffusion-based text-to-video generation. VidCLearn features a student-teacher architecture where the student model is incrementally updated with new text-video pairs, and the teacher model helps preserve previously learned knowledge through generative replay. Additionally, we introduce a novel temporal consistency loss to enhance motion smoothness and a video retrieval module to provide structural guidance at inference. Our architecture is also designed to be more computationally efficient than existing models while retaining satisfactory generation performance. Experimental results show VidCLearn's superiority over baseline methods in terms of visual quality, semantic alignment, and temporal coherence.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16956
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VidCLearn: A Continual Learning Approach for Text-to-Video Generation
Zanchetta, Luca
Papa, Lorenzo
Maiano, Luca
Amerini, Irene
Computer Vision and Pattern Recognition
Text-to-video generation is an emerging field in generative AI, enabling the creation of realistic, semantically accurate videos from text prompts. While current models achieve impressive visual quality and alignment with input text, they typically rely on static knowledge, making it difficult to incorporate new data without retraining from scratch. To address this limitation, we propose VidCLearn, a continual learning framework for diffusion-based text-to-video generation. VidCLearn features a student-teacher architecture where the student model is incrementally updated with new text-video pairs, and the teacher model helps preserve previously learned knowledge through generative replay. Additionally, we introduce a novel temporal consistency loss to enhance motion smoothness and a video retrieval module to provide structural guidance at inference. Our architecture is also designed to be more computationally efficient than existing models while retaining satisfactory generation performance. Experimental results show VidCLearn's superiority over baseline methods in terms of visual quality, semantic alignment, and temporal coherence.
title VidCLearn: A Continual Learning Approach for Text-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.16956