InstructVid2Vid: Controllable Video Editing with Natural Language Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Bosheng, Li, Juncheng, Tang, Siliang, Chua, Tat-Seng, Zhuang, Yueting
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929362751717376
author Qin, Bosheng
Li, Juncheng
Tang, Siliang
Chua, Tat-Seng
Zhuang, Yueting
author_facet Qin, Bosheng
Li, Juncheng
Tang, Siliang
Chua, Tat-Seng
Zhuang, Yueting
contents We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating the need for per-example fine-tuning or inversion. The proposed InstructVid2Vid model modifies a pretrained image generation model, Stable Diffusion, to generate a time-dependent sequence of video frames. By harnessing the collective intelligence of disparate models, we engineer a training dataset rich in video-instruction triplets, which is a more cost-efficient alternative to collecting data in real-world scenarios. To enhance the coherence between successive frames within the generated videos, we propose the Inter-Frames Consistency Loss and incorporate it during the training process. With multimodal classifier-free guidance during the inference stage, the generated videos is able to resonate with both the input video and the accompanying instructions. Experimental results demonstrate that InstructVid2Vid is capable of generating high-quality, temporally coherent videos and performing diverse edits, including attribute editing, background changes, and style transfer. These results underscore the versatility and effectiveness of our proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2305_12328
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
Qin, Bosheng
Li, Juncheng
Tang, Siliang
Chua, Tat-Seng
Zhuang, Yueting
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating the need for per-example fine-tuning or inversion. The proposed InstructVid2Vid model modifies a pretrained image generation model, Stable Diffusion, to generate a time-dependent sequence of video frames. By harnessing the collective intelligence of disparate models, we engineer a training dataset rich in video-instruction triplets, which is a more cost-efficient alternative to collecting data in real-world scenarios. To enhance the coherence between successive frames within the generated videos, we propose the Inter-Frames Consistency Loss and incorporate it during the training process. With multimodal classifier-free guidance during the inference stage, the generated videos is able to resonate with both the input video and the accompanying instructions. Experimental results demonstrate that InstructVid2Vid is capable of generating high-quality, temporally coherent videos and performing diverse edits, including attribute editing, background changes, and style transfer. These results underscore the versatility and effectiveness of our proposed method.
title InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2305.12328