Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Wen, Jiang, Yan, Xie, Kangyang, Liu, Zide, Chen, Hao, Cao, Yue, Wang, Xinlong, Shen, Chunhua
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911746954887168
author Wang, Wen
Jiang, Yan
Xie, Kangyang
Liu, Zide
Chen, Hao
Cao, Yue
Wang, Xinlong
Shen, Chunhua
author_facet Wang, Wen
Jiang, Yan
Xie, Kangyang
Liu, Zide
Chen, Hao
Cao, Yue
Wang, Xinlong
Shen, Chunhua
contents Large-scale text-to-image diffusion models achieve unprecedented success in image generation and editing. However, how to extend such success to video editing is unclear. Recent initial attempts at video editing require significant text-to-video data and computation resources for training, which is often not accessible. In this work, we propose vid2vid-zero, a simple yet effective method for zero-shot video editing. Our vid2vid-zero leverages off-the-shelf image diffusion models, and doesn't require training on any video. At the core of our method is a null-text inversion module for text-to-video alignment, a cross-frame modeling module for temporal consistency, and a spatial regularization module for fidelity to the original video. Without any training, we leverage the dynamic nature of the attention mechanism to enable bi-directional temporal modeling at test time. Experiments and analyses show promising results in editing attributes, subjects, places, etc., in real-world videos. Code is made available at \url{https://github.com/baaivision/vid2vid-zero}.
format Preprint
id arxiv_https___arxiv_org_abs_2303_17599
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models
Wang, Wen
Jiang, Yan
Xie, Kangyang
Liu, Zide
Chen, Hao
Cao, Yue
Wang, Xinlong
Shen, Chunhua
Computer Vision and Pattern Recognition
Large-scale text-to-image diffusion models achieve unprecedented success in image generation and editing. However, how to extend such success to video editing is unclear. Recent initial attempts at video editing require significant text-to-video data and computation resources for training, which is often not accessible. In this work, we propose vid2vid-zero, a simple yet effective method for zero-shot video editing. Our vid2vid-zero leverages off-the-shelf image diffusion models, and doesn't require training on any video. At the core of our method is a null-text inversion module for text-to-video alignment, a cross-frame modeling module for temporal consistency, and a spatial regularization module for fidelity to the original video. Without any training, we leverage the dynamic nature of the attention mechanism to enable bi-directional temporal modeling at test time. Experiments and analyses show promising results in editing attributes, subjects, places, etc., in real-world videos. Code is made available at \url{https://github.com/baaivision/vid2vid-zero}.
title Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.17599