VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Dohun, Kim, Bryan S, Park, Geon Yeong, Ye, Jong Chul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916512144556032
author Lee, Dohun
Kim, Bryan S
Park, Geon Yeong
Ye, Jong Chul
author_facet Lee, Dohun
Kim, Bryan S
Park, Geon Yeong
Ye, Jong Chul
contents Text-to-image (T2I) diffusion models have revolutionized visual content creation, but extending these capabilities to text-to-video (T2V) generation remains a challenge, particularly in preserving temporal consistency. Existing methods that aim to improve consistency often cause trade-offs such as reduced imaging quality and impractical computational time. To address these issues we introduce VideoGuide, a novel framework that enhances the temporal consistency of pretrained T2V models without the need for additional training or fine-tuning. Instead, VideoGuide leverages any pretrained video diffusion model (VDM) or itself as a guide during the early stages of inference, improving temporal quality by interpolating the guiding model's denoised samples into the sampling model's denoising process. The proposed method brings about significant improvement in temporal consistency and image fidelity, providing a cost-effective and practical solution that synergizes the strengths of various video diffusion models. Furthermore, we demonstrate prior distillation, revealing that base models can achieve enhanced text coherence by utilizing the superior data prior of the guiding model through the proposed method. Project Page: https://dohunlee1.github.io/videoguide.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2410_04364
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide
Lee, Dohun
Kim, Bryan S
Park, Geon Yeong
Ye, Jong Chul
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Text-to-image (T2I) diffusion models have revolutionized visual content creation, but extending these capabilities to text-to-video (T2V) generation remains a challenge, particularly in preserving temporal consistency. Existing methods that aim to improve consistency often cause trade-offs such as reduced imaging quality and impractical computational time. To address these issues we introduce VideoGuide, a novel framework that enhances the temporal consistency of pretrained T2V models without the need for additional training or fine-tuning. Instead, VideoGuide leverages any pretrained video diffusion model (VDM) or itself as a guide during the early stages of inference, improving temporal quality by interpolating the guiding model's denoised samples into the sampling model's denoising process. The proposed method brings about significant improvement in temporal consistency and image fidelity, providing a cost-effective and practical solution that synergizes the strengths of various video diffusion models. Furthermore, we demonstrate prior distillation, revealing that base models can achieve enhanced text coherence by utilizing the superior data prior of the guiding model through the proposed method. Project Page: https://dohunlee1.github.io/videoguide.github.io/
title VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.04364