InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Saini, Nirat, Bodla, Navaneeth, Shrivastava, Ashish, Ravichandran, Avinash, Zhang, Xiao, Shrivastava, Abhinav, Singh, Bharat
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916324790239232
author Saini, Nirat
Bodla, Navaneeth
Shrivastava, Ashish
Ravichandran, Avinash
Zhang, Xiao
Shrivastava, Abhinav
Singh, Bharat
author_facet Saini, Nirat
Bodla, Navaneeth
Shrivastava, Ashish
Ravichandran, Avinash
Zhang, Xiao
Shrivastava, Abhinav
Singh, Bharat
contents We introduce InVi, an approach for inserting or replacing objects within videos (referred to as inpainting) using off-the-shelf, text-to-image latent diffusion models. InVi targets controlled manipulation of objects and blending them seamlessly into a background video unlike existing video editing methods that focus on comprehensive re-styling or entire scene alterations. To achieve this goal, we tackle two key challenges. Firstly, for high quality control and blending, we employ a two-step process involving inpainting and matching. This process begins with inserting the object into a single frame using a ControlNet-based inpainting diffusion model, and then generating subsequent frames conditioned on features from an inpainted frame as an anchor to minimize the domain gap between the background and the object. Secondly, to ensure temporal coherence, we replace the diffusion model's self-attention layers with extended-attention layers. The anchor frame features serve as the keys and values for these layers, enhancing consistency across frames. Our approach removes the need for video-specific fine-tuning, presenting an efficient and adaptable solution. Experimental results demonstrate that InVi achieves realistic object insertion with consistent blending and coherence across frames, outperforming existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10958
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
Saini, Nirat
Bodla, Navaneeth
Shrivastava, Ashish
Ravichandran, Avinash
Zhang, Xiao
Shrivastava, Abhinav
Singh, Bharat
Computer Vision and Pattern Recognition
We introduce InVi, an approach for inserting or replacing objects within videos (referred to as inpainting) using off-the-shelf, text-to-image latent diffusion models. InVi targets controlled manipulation of objects and blending them seamlessly into a background video unlike existing video editing methods that focus on comprehensive re-styling or entire scene alterations. To achieve this goal, we tackle two key challenges. Firstly, for high quality control and blending, we employ a two-step process involving inpainting and matching. This process begins with inserting the object into a single frame using a ControlNet-based inpainting diffusion model, and then generating subsequent frames conditioned on features from an inpainted frame as an anchor to minimize the domain gap between the background and the object. Secondly, to ensure temporal coherence, we replace the diffusion model's self-attention layers with extended-attention layers. The anchor frame features serve as the keys and values for these layers, enhancing consistency across frames. Our approach removes the need for video-specific fine-tuning, presenting an efficient and adaptable solution. Experimental results demonstrate that InVi achieves realistic object insertion with consistent blending and coherence across frames, outperforming existing methods.
title InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.10958