Training-Free Semantic Video Composition via Pre-trained Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Jiaqi, Su, Sitong, Zhu, Junchen, Gao, Lianli, Song, Jingkuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911759675162624
author Guo, Jiaqi
Su, Sitong
Zhu, Junchen
Gao, Lianli
Song, Jingkuan
author_facet Guo, Jiaqi
Su, Sitong
Zhu, Junchen
Gao, Lianli
Song, Jingkuan
contents The video composition task aims to integrate specified foregrounds and backgrounds from different videos into a harmonious composite. Current approaches, predominantly trained on videos with adjusted foreground color and lighting, struggle to address deep semantic disparities beyond superficial adjustments, such as domain gaps. Therefore, we propose a training-free pipeline employing a pre-trained diffusion model imbued with semantic prior knowledge, which can process composite videos with broader semantic disparities. Specifically, we process the video frames in a cascading manner and handle each frame in two processes with the diffusion model. In the inversion process, we propose Balanced Partial Inversion to obtain generation initial points that balance reversibility and modifiability. Then, in the generation process, we further propose Inter-Frame Augmented attention to augment foreground continuity across frames. Experimental results reveal that our pipeline successfully ensures the visual harmony and inter-frame coherence of the outputs, demonstrating efficacy in managing broader semantic disparities.
format Preprint
id arxiv_https___arxiv_org_abs_2401_09195
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Training-Free Semantic Video Composition via Pre-trained Diffusion Model
Guo, Jiaqi
Su, Sitong
Zhu, Junchen
Gao, Lianli
Song, Jingkuan
Computer Vision and Pattern Recognition
The video composition task aims to integrate specified foregrounds and backgrounds from different videos into a harmonious composite. Current approaches, predominantly trained on videos with adjusted foreground color and lighting, struggle to address deep semantic disparities beyond superficial adjustments, such as domain gaps. Therefore, we propose a training-free pipeline employing a pre-trained diffusion model imbued with semantic prior knowledge, which can process composite videos with broader semantic disparities. Specifically, we process the video frames in a cascading manner and handle each frame in two processes with the diffusion model. In the inversion process, we propose Balanced Partial Inversion to obtain generation initial points that balance reversibility and modifiability. Then, in the generation process, we further propose Inter-Frame Augmented attention to augment foreground continuity across frames. Experimental results reveal that our pipeline successfully ensures the visual harmony and inter-frame coherence of the outputs, demonstrating efficacy in managing broader semantic disparities.
title Training-Free Semantic Video Composition via Pre-trained Diffusion Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.09195