Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Tianyi, Zhang, Xing, Gu, Jiaxi, Pei, Renjing, Xu, Songcen, Ma, Xingjun, Xu, Hang, Wu, Zuxuan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916426445488128
author Lu, Tianyi
Zhang, Xing
Gu, Jiaxi
Pei, Renjing
Xu, Songcen
Ma, Xingjun
Xu, Hang
Wu, Zuxuan
author_facet Lu, Tianyi
Zhang, Xing
Gu, Jiaxi
Pei, Renjing
Xu, Songcen
Ma, Xingjun
Xu, Hang
Wu, Zuxuan
contents Latent Diffusion Models (LDMs) are renowned for their powerful capabilities in image and video synthesis. Yet, compared to text-to-image (T2I) editing, text-to-video (T2V) editing suffers from a lack of decent temporal consistency and structure, due to insufficient pre-training data, limited model editability, or extensive tuning costs. To address this gap, we propose FLDM (Fused Latent Diffusion Model), a training-free framework that achieves high-quality T2V editing by integrating various T2I and T2V LDMs. Specifically, FLDM utilizes a hyper-parameter with an update schedule to effectively fuse image and video latents during the denoising process. This paper is the first to reveal that T2I and T2V LDMs can complement each other in terms of structure and temporal consistency, ultimately generating high-quality videos. It is worth noting that FLDM can serve as a versatile plugin, applicable to off-the-shelf image and video LDMs, to significantly enhance the quality of video editing. Extensive quantitative and qualitative experiments on popular T2I and T2V LDMs demonstrate FLDM's superior editing quality than state-of-the-art T2V editing methods. Our project code is available at https://github.com/lutianyi0603/fuse_your_latents.
format Preprint
id arxiv_https___arxiv_org_abs_2310_16400
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
Lu, Tianyi
Zhang, Xing
Gu, Jiaxi
Pei, Renjing
Xu, Songcen
Ma, Xingjun
Xu, Hang
Wu, Zuxuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Latent Diffusion Models (LDMs) are renowned for their powerful capabilities in image and video synthesis. Yet, compared to text-to-image (T2I) editing, text-to-video (T2V) editing suffers from a lack of decent temporal consistency and structure, due to insufficient pre-training data, limited model editability, or extensive tuning costs. To address this gap, we propose FLDM (Fused Latent Diffusion Model), a training-free framework that achieves high-quality T2V editing by integrating various T2I and T2V LDMs. Specifically, FLDM utilizes a hyper-parameter with an update schedule to effectively fuse image and video latents during the denoising process. This paper is the first to reveal that T2I and T2V LDMs can complement each other in terms of structure and temporal consistency, ultimately generating high-quality videos. It is worth noting that FLDM can serve as a versatile plugin, applicable to off-the-shelf image and video LDMs, to significantly enhance the quality of video editing. Extensive quantitative and qualitative experiments on popular T2I and T2V LDMs demonstrate FLDM's superior editing quality than state-of-the-art T2V editing methods. Our project code is available at https://github.com/lutianyi0603/fuse_your_latents.
title Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2310.16400