ObjectDrop: Bootstrapping Counterfactuals for Photorealistic Object Removal and Insertion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Winter, Daniel, Cohen, Matan, Fruchter, Shlomi, Pritch, Yael, Rav-Acha, Alex, Hoshen, Yedid
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913286914572288
author Winter, Daniel
Cohen, Matan
Fruchter, Shlomi
Pritch, Yael
Rav-Acha, Alex
Hoshen, Yedid
author_facet Winter, Daniel
Cohen, Matan
Fruchter, Shlomi
Pritch, Yael
Rav-Acha, Alex
Hoshen, Yedid
contents Diffusion models have revolutionized image editing but often generate images that violate physical laws, particularly the effects of objects on the scene, e.g., occlusions, shadows, and reflections. By analyzing the limitations of self-supervised approaches, we propose a practical solution centered on a \q{counterfactual} dataset. Our method involves capturing a scene before and after removing a single object, while minimizing other changes. By fine-tuning a diffusion model on this dataset, we are able to not only remove objects but also their effects on the scene. However, we find that applying this approach for photorealistic object insertion requires an impractically large dataset. To tackle this challenge, we propose bootstrap supervision; leveraging our object removal model trained on a small counterfactual dataset, we synthetically expand this dataset considerably. Our approach significantly outperforms prior methods in photorealistic object removal and insertion, particularly at modeling the effects of objects on the scene.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18818
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ObjectDrop: Bootstrapping Counterfactuals for Photorealistic Object Removal and Insertion
Winter, Daniel
Cohen, Matan
Fruchter, Shlomi
Pritch, Yael
Rav-Acha, Alex
Hoshen, Yedid
Computer Vision and Pattern Recognition
Diffusion models have revolutionized image editing but often generate images that violate physical laws, particularly the effects of objects on the scene, e.g., occlusions, shadows, and reflections. By analyzing the limitations of self-supervised approaches, we propose a practical solution centered on a \q{counterfactual} dataset. Our method involves capturing a scene before and after removing a single object, while minimizing other changes. By fine-tuning a diffusion model on this dataset, we are able to not only remove objects but also their effects on the scene. However, we find that applying this approach for photorealistic object insertion requires an impractically large dataset. To tackle this challenge, we propose bootstrap supervision; leveraging our object removal model trained on a small counterfactual dataset, we synthetically expand this dataset considerably. Our approach significantly outperforms prior methods in photorealistic object removal and insertion, particularly at modeling the effects of objects on the scene.
title ObjectDrop: Bootstrapping Counterfactuals for Photorealistic Object Removal and Insertion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.18818