TVG: A Training-free Transition Video Generation Method with Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Rui, Chen, Yaosen, Liu, Yuegen, Wang, Wei, Wen, Xuming, Wang, Hongxia
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916368817848320
author Zhang, Rui
Chen, Yaosen
Liu, Yuegen
Wang, Wei
Wen, Xuming
Wang, Hongxia
author_facet Zhang, Rui
Chen, Yaosen
Liu, Yuegen
Wang, Wei
Wen, Xuming
Wang, Hongxia
contents Transition videos play a crucial role in media production, enhancing the flow and coherence of visual narratives. Traditional methods like morphing often lack artistic appeal and require specialized skills, limiting their effectiveness. Recent advances in diffusion model-based video generation offer new possibilities for creating transitions but face challenges such as poor inter-frame relationship modeling and abrupt content changes. We propose a novel training-free Transition Video Generation (TVG) approach using video-level diffusion models that addresses these limitations without additional training. Our method leverages Gaussian Process Regression ($\mathcal{GPR}$) to model latent representations, ensuring smooth and dynamic transitions between frames. Additionally, we introduce interpolation-based conditional controls and a Frequency-aware Bidirectional Fusion (FBiF) architecture to enhance temporal control and transition reliability. Evaluations of benchmark datasets and custom image pairs demonstrate the effectiveness of our approach in generating high-quality smooth transition videos. The code are provided in https://sobeymil.github.io/tvg.com.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13413
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TVG: A Training-free Transition Video Generation Method with Diffusion Models
Zhang, Rui
Chen, Yaosen
Liu, Yuegen
Wang, Wei
Wen, Xuming
Wang, Hongxia
Computer Vision and Pattern Recognition
Transition videos play a crucial role in media production, enhancing the flow and coherence of visual narratives. Traditional methods like morphing often lack artistic appeal and require specialized skills, limiting their effectiveness. Recent advances in diffusion model-based video generation offer new possibilities for creating transitions but face challenges such as poor inter-frame relationship modeling and abrupt content changes. We propose a novel training-free Transition Video Generation (TVG) approach using video-level diffusion models that addresses these limitations without additional training. Our method leverages Gaussian Process Regression ($\mathcal{GPR}$) to model latent representations, ensuring smooth and dynamic transitions between frames. Additionally, we introduce interpolation-based conditional controls and a Frequency-aware Bidirectional Fusion (FBiF) architecture to enhance temporal control and transition reliability. Evaluations of benchmark datasets and custom image pairs demonstrate the effectiveness of our approach in generating high-quality smooth transition videos. The code are provided in https://sobeymil.github.io/tvg.com.
title TVG: A Training-free Transition Video Generation Method with Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.13413