StableV2V: Stablizing Shape Consistency in Video-to-Video Editing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Chang, Li, Rui, Zhang, Kaidong, Lan, Yunwei, Liu, Dong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915649596424192
author Liu, Chang
Li, Rui
Zhang, Kaidong
Lan, Yunwei
Liu, Dong
author_facet Liu, Chang
Li, Rui
Zhang, Kaidong
Lan, Yunwei
Liu, Dong
contents Recent advancements of generative AI have significantly promoted content creation and editing, where prevailing studies further extend this exciting progress to video editing. In doing so, these studies mainly transfer the inherent motion patterns from the source videos to the edited ones, where results with inferior consistency to user prompts are often observed, due to the lack of particular alignments between the delivered motions and edited contents. To address this limitation, we present a shape-consistent video editing method, namely StableV2V, in this paper. Our method decomposes the entire editing pipeline into several sequential procedures, where it edits the first video frame, then establishes an alignment between the delivered motions and user prompts, and eventually propagates the edited contents to all other frames based on such alignment. Furthermore, we curate a testing benchmark, namely DAVIS-Edit, for a comprehensive evaluation of video editing, considering various types of prompts and difficulties. Experimental results and analyses illustrate the outperforming performance, visual consistency, and inference efficiency of our method compared to existing state-of-the-art studies.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11045
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StableV2V: Stablizing Shape Consistency in Video-to-Video Editing
Liu, Chang
Li, Rui
Zhang, Kaidong
Lan, Yunwei
Liu, Dong
Computer Vision and Pattern Recognition
Recent advancements of generative AI have significantly promoted content creation and editing, where prevailing studies further extend this exciting progress to video editing. In doing so, these studies mainly transfer the inherent motion patterns from the source videos to the edited ones, where results with inferior consistency to user prompts are often observed, due to the lack of particular alignments between the delivered motions and edited contents. To address this limitation, we present a shape-consistent video editing method, namely StableV2V, in this paper. Our method decomposes the entire editing pipeline into several sequential procedures, where it edits the first video frame, then establishes an alignment between the delivered motions and user prompts, and eventually propagates the edited contents to all other frames based on such alignment. Furthermore, we curate a testing benchmark, namely DAVIS-Edit, for a comprehensive evaluation of video editing, considering various types of prompts and difficulties. Experimental results and analyses illustrate the outperforming performance, visual consistency, and inference efficiency of our method compared to existing state-of-the-art studies.
title StableV2V: Stablizing Shape Consistency in Video-to-Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.11045