VideoDirector: Precise Video Editing via Text-to-Video Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yukun, Wang, Longguang, Ma, Zhiyuan, Hu, Qibin, Xu, Kai, Guo, Yulan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909541825773568
author Wang, Yukun
Wang, Longguang
Ma, Zhiyuan
Hu, Qibin
Xu, Kai
Guo, Yulan
author_facet Wang, Yukun
Wang, Longguang
Ma, Zhiyuan
Hu, Qibin
Xu, Kai
Guo, Yulan
contents Despite the typical inversion-then-editing paradigm using text-to-image (T2I) models has demonstrated promising results, directly extending it to text-to-video (T2V) models still suffers severe artifacts such as color flickering and content distortion. Consequently, current video editing methods primarily rely on T2I models, which inherently lack temporal-coherence generative ability, often resulting in inferior editing results. In this paper, we attribute the failure of the typical editing paradigm to: 1) Tightly Spatial-temporal Coupling. The vanilla pivotal-based inversion strategy struggles to disentangle spatial-temporal information in the video diffusion model; 2) Complicated Spatial-temporal Layout. The vanilla cross-attention control is deficient in preserving the unedited content. To address these limitations, we propose a spatial-temporal decoupled guidance (STDG) and multi-frame null-text optimization strategy to provide pivotal temporal cues for more precise pivotal inversion. Furthermore, we introduce a self-attention control strategy to maintain higher fidelity for precise partial content editing. Experimental results demonstrate that our method (termed VideoDirector) effectively harnesses the powerful temporal generation capabilities of T2V models, producing edited videos with state-of-the-art performance in accuracy, motion smoothness, realism, and fidelity to unedited content.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17592
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VideoDirector: Precise Video Editing via Text-to-Video Models
Wang, Yukun
Wang, Longguang
Ma, Zhiyuan
Hu, Qibin
Xu, Kai
Guo, Yulan
Computer Vision and Pattern Recognition
Despite the typical inversion-then-editing paradigm using text-to-image (T2I) models has demonstrated promising results, directly extending it to text-to-video (T2V) models still suffers severe artifacts such as color flickering and content distortion. Consequently, current video editing methods primarily rely on T2I models, which inherently lack temporal-coherence generative ability, often resulting in inferior editing results. In this paper, we attribute the failure of the typical editing paradigm to: 1) Tightly Spatial-temporal Coupling. The vanilla pivotal-based inversion strategy struggles to disentangle spatial-temporal information in the video diffusion model; 2) Complicated Spatial-temporal Layout. The vanilla cross-attention control is deficient in preserving the unedited content. To address these limitations, we propose a spatial-temporal decoupled guidance (STDG) and multi-frame null-text optimization strategy to provide pivotal temporal cues for more precise pivotal inversion. Furthermore, we introduce a self-attention control strategy to maintain higher fidelity for precise partial content editing. Experimental results demonstrate that our method (termed VideoDirector) effectively harnesses the powerful temporal generation capabilities of T2V models, producing edited videos with state-of-the-art performance in accuracy, motion smoothness, realism, and fidelity to unedited content.
title VideoDirector: Precise Video Editing via Text-to-Video Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.17592