Edit Temporal-Consistent Videos with Image Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yuanzhi, Li, Yong, Zhang, Xiaoya, Liu, Xin, Dai, Anbo, Chan, Antoni B., Cui, Zhen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914926023409664
author Wang, Yuanzhi
Li, Yong
Zhang, Xiaoya
Liu, Xin
Dai, Anbo
Chan, Antoni B.
Cui, Zhen
author_facet Wang, Yuanzhi
Li, Yong
Zhang, Xiaoya
Liu, Xin
Dai, Anbo
Chan, Antoni B.
Cui, Zhen
contents Large-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregularities and temporal inconsistencies as the temporal characteristics of videos have not been faithfully modeled. In this paper, we propose an elegant yet effective Temporal-Consistent Video Editing (TCVE) method to mitigate the temporal inconsistency challenge for robust text-guided video editing. In addition to the utilization of a pretrained T2I 2D Unet for spatial content manipulation, we establish a dedicated temporal Unet architecture to faithfully capture the temporal coherence of the input video sequences. Furthermore, to establish coherence and interrelation between the spatial-focused and temporal-focused components, a cohesive spatial-temporal modeling unit is formulated. This unit effectively interconnects the temporal Unet with the pretrained 2D Unet, thereby enhancing the temporal consistency of the generated videos while preserving the capacity for video content manipulation. Quantitative experimental results and visualization results demonstrate that TCVE achieves state-of-the-art performance in both video temporal consistency and video editing capability, surpassing existing benchmarks in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2308_09091
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Edit Temporal-Consistent Videos with Image Diffusion Model
Wang, Yuanzhi
Li, Yong
Zhang, Xiaoya
Liu, Xin
Dai, Anbo
Chan, Antoni B.
Cui, Zhen
Computer Vision and Pattern Recognition
Large-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregularities and temporal inconsistencies as the temporal characteristics of videos have not been faithfully modeled. In this paper, we propose an elegant yet effective Temporal-Consistent Video Editing (TCVE) method to mitigate the temporal inconsistency challenge for robust text-guided video editing. In addition to the utilization of a pretrained T2I 2D Unet for spatial content manipulation, we establish a dedicated temporal Unet architecture to faithfully capture the temporal coherence of the input video sequences. Furthermore, to establish coherence and interrelation between the spatial-focused and temporal-focused components, a cohesive spatial-temporal modeling unit is formulated. This unit effectively interconnects the temporal Unet with the pretrained 2D Unet, thereby enhancing the temporal consistency of the generated videos while preserving the capacity for video content manipulation. Quantitative experimental results and visualization results demonstrate that TCVE achieves state-of-the-art performance in both video temporal consistency and video editing capability, surpassing existing benchmarks in the field.
title Edit Temporal-Consistent Videos with Image Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.09091