Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Su, Tongtong, Wang, Chengyu, Liu, Bingyan, Huang, Jun, Lu, Dongming
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912490608132096
author Su, Tongtong
Wang, Chengyu
Liu, Bingyan
Huang, Jun
Lu, Dongming
author_facet Su, Tongtong
Wang, Chengyu
Liu, Bingyan
Huang, Jun
Lu, Dongming
contents In recent years, large text-to-video (T2V) synthesis models have garnered considerable attention for their abilities to generate videos from textual descriptions. However, achieving both high imaging quality and effective motion representation remains a significant challenge for these T2V models. Existing approaches often adapt pre-trained text-to-image (T2I) models to refine video frames, leading to issues such as flickering and artifacts due to inconsistencies across frames. In this paper, we introduce EVS, a training-free Encapsulated Video Synthesizer that composes T2I and T2V models to enhance both visual fidelity and motion smoothness of generated videos. Our approach utilizes a well-trained diffusion-based T2I model to refine low-quality video frames by treating them as out-of-distribution samples, effectively optimizing them with noising and denoising steps. Meanwhile, we employ T2V backbones to ensure consistent motion dynamics. By encapsulating the T2V temporal-only prior into the T2I generation process, EVS successfully leverages the strengths of both types of models, resulting in videos of improved imaging and motion quality. Experimental results validate the effectiveness of our approach compared to previous approaches. Our composition process also leads to a significant improvement of 1.6x-4.5x speedup in inference time. Source codes: https://github.com/Tonniia/EVS.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
Su, Tongtong
Wang, Chengyu
Liu, Bingyan
Huang, Jun
Lu, Dongming
Computer Vision and Pattern Recognition
In recent years, large text-to-video (T2V) synthesis models have garnered considerable attention for their abilities to generate videos from textual descriptions. However, achieving both high imaging quality and effective motion representation remains a significant challenge for these T2V models. Existing approaches often adapt pre-trained text-to-image (T2I) models to refine video frames, leading to issues such as flickering and artifacts due to inconsistencies across frames. In this paper, we introduce EVS, a training-free Encapsulated Video Synthesizer that composes T2I and T2V models to enhance both visual fidelity and motion smoothness of generated videos. Our approach utilizes a well-trained diffusion-based T2I model to refine low-quality video frames by treating them as out-of-distribution samples, effectively optimizing them with noising and denoising steps. Meanwhile, we employ T2V backbones to ensure consistent motion dynamics. By encapsulating the T2V temporal-only prior into the T2I generation process, EVS successfully leverages the strengths of both types of models, resulting in videos of improved imaging and motion quality. Experimental results validate the effectiveness of our approach compared to previous approaches. Our composition process also leads to a significant improvement of 1.6x-4.5x speedup in inference time. Source codes: https://github.com/Tonniia/EVS.
title Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.13753