Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Mingce, He, Jingxuan, Tang, Shengeng, Wang, Zhangye, Cheng, Lechao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909623927177216
author Guo, Mingce
He, Jingxuan
Tang, Shengeng
Wang, Zhangye
Cheng, Lechao
author_facet Guo, Mingce
He, Jingxuan
Tang, Shengeng
Wang, Zhangye
Cheng, Lechao
contents Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided in pre-training, which hinders nuanced editing targeting open concepts with specific attributes. Directly altering the keywords in target prompts often results in unintended disruptions to the attention mechanisms. To achieve more flexible editing easily, this work proposes an improved concept-augmented video editing approach that generates diverse and stable target videos flexibly by devising abstract conceptual pairs. Specifically, the framework involves concept-augmented textual inversion and a dual prior supervision mechanism. The former enables plug-and-play guidance of stable diffusion for video editing, effectively capturing target attributes for more stylized results. The dual prior supervision mechanism significantly enhances video stability and fidelity. Comprehensive evaluations demonstrate that our approach generates more stable and lifelike videos, outperforming state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12526
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing
Guo, Mingce
He, Jingxuan
Tang, Shengeng
Wang, Zhangye
Cheng, Lechao
Computer Vision and Pattern Recognition
Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided in pre-training, which hinders nuanced editing targeting open concepts with specific attributes. Directly altering the keywords in target prompts often results in unintended disruptions to the attention mechanisms. To achieve more flexible editing easily, this work proposes an improved concept-augmented video editing approach that generates diverse and stable target videos flexibly by devising abstract conceptual pairs. Specifically, the framework involves concept-augmented textual inversion and a dual prior supervision mechanism. The former enables plug-and-play guidance of stable diffusion for video editing, effectively capturing target attributes for more stylized results. The dual prior supervision mechanism significantly enhances video stability and fidelity. Comprehensive evaluations demonstrate that our approach generates more stable and lifelike videos, outperforming state-of-the-art methods.
title Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.12526