Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yuyang, Xie, Enze, Hong, Lanqing, Li, Zhenguo, Lee, Gim Hee
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911779180773376
author Zhao, Yuyang
Xie, Enze
Hong, Lanqing
Li, Zhenguo
Lee, Gim Hee
author_facet Zhao, Yuyang
Xie, Enze
Hong, Lanqing
Li, Zhenguo
Lee, Gim Hee
contents The text-driven image and video diffusion models have achieved unprecedented success in generating realistic and diverse content. Recently, the editing and variation of existing images and videos in diffusion-based generative models have garnered significant attention. However, previous works are limited to editing content with text or providing coarse personalization using a single visual clue, rendering them unsuitable for indescribable content that requires fine-grained and detailed control. In this regard, we propose a generic video editing framework called Make-A-Protagonist, which utilizes textual and visual clues to edit videos with the goal of empowering individuals to become the protagonists. Specifically, we leverage multiple experts to parse source video, target visual and textual clues, and propose a visual-textual-based video generation model that employs mask-guided denoising sampling to generate the desired output. Extensive results demonstrate the versatile and remarkable editing capabilities of Make-A-Protagonist.
format Preprint
id arxiv_https___arxiv_org_abs_2305_08850
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
Zhao, Yuyang
Xie, Enze
Hong, Lanqing
Li, Zhenguo
Lee, Gim Hee
Computer Vision and Pattern Recognition
The text-driven image and video diffusion models have achieved unprecedented success in generating realistic and diverse content. Recently, the editing and variation of existing images and videos in diffusion-based generative models have garnered significant attention. However, previous works are limited to editing content with text or providing coarse personalization using a single visual clue, rendering them unsuitable for indescribable content that requires fine-grained and detailed control. In this regard, we propose a generic video editing framework called Make-A-Protagonist, which utilizes textual and visual clues to edit videos with the goal of empowering individuals to become the protagonists. Specifically, we leverage multiple experts to parse source video, target visual and textual clues, and propose a visual-textual-based video generation model that employs mask-guided denoising sampling to generate the desired output. Extensive results demonstrate the versatile and remarkable editing capabilities of Make-A-Protagonist.
title Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.08850