RefFusion: Reference Adapted Diffusion Models for 3D Scene Inpainting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mirzaei, Ashkan, De Lutio, Riccardo, Kim, Seung Wook, Acuna, David, Kelly, Jonathan, Fidler, Sanja, Gilitschenski, Igor, Gojcic, Zan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914757234130944
author Mirzaei, Ashkan
De Lutio, Riccardo
Kim, Seung Wook
Acuna, David
Kelly, Jonathan
Fidler, Sanja
Gilitschenski, Igor
Gojcic, Zan
author_facet Mirzaei, Ashkan
De Lutio, Riccardo
Kim, Seung Wook
Acuna, David
Kelly, Jonathan
Fidler, Sanja
Gilitschenski, Igor
Gojcic, Zan
contents Neural reconstruction approaches are rapidly emerging as the preferred representation for 3D scenes, but their limited editability is still posing a challenge. In this work, we propose an approach for 3D scene inpainting -- the task of coherently replacing parts of the reconstructed scene with desired content. Scene inpainting is an inherently ill-posed task as there exist many solutions that plausibly replace the missing content. A good inpainting method should therefore not only enable high-quality synthesis but also a high degree of control. Based on this observation, we focus on enabling explicit control over the inpainted content and leverage a reference image as an efficient means to achieve this goal. Specifically, we introduce RefFusion, a novel 3D inpainting method based on a multi-scale personalization of an image inpainting diffusion model to the given reference view. The personalization effectively adapts the prior distribution to the target scene, resulting in a lower variance of score distillation objective and hence significantly sharper details. Our framework achieves state-of-the-art results for object removal while maintaining high controllability. We further demonstrate the generality of our formulation on other downstream tasks such as object insertion, scene outpainting, and sparse view reconstruction.
format Preprint
id arxiv_https___arxiv_org_abs_2404_10765
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RefFusion: Reference Adapted Diffusion Models for 3D Scene Inpainting
Mirzaei, Ashkan
De Lutio, Riccardo
Kim, Seung Wook
Acuna, David
Kelly, Jonathan
Fidler, Sanja
Gilitschenski, Igor
Gojcic, Zan
Computer Vision and Pattern Recognition
Neural reconstruction approaches are rapidly emerging as the preferred representation for 3D scenes, but their limited editability is still posing a challenge. In this work, we propose an approach for 3D scene inpainting -- the task of coherently replacing parts of the reconstructed scene with desired content. Scene inpainting is an inherently ill-posed task as there exist many solutions that plausibly replace the missing content. A good inpainting method should therefore not only enable high-quality synthesis but also a high degree of control. Based on this observation, we focus on enabling explicit control over the inpainted content and leverage a reference image as an efficient means to achieve this goal. Specifically, we introduce RefFusion, a novel 3D inpainting method based on a multi-scale personalization of an image inpainting diffusion model to the given reference view. The personalization effectively adapts the prior distribution to the target scene, resulting in a lower variance of score distillation objective and hence significantly sharper details. Our framework achieves state-of-the-art results for object removal while maintaining high controllability. We further demonstrate the generality of our formulation on other downstream tasks such as object insertion, scene outpainting, and sparse view reconstruction.
title RefFusion: Reference Adapted Diffusion Models for 3D Scene Inpainting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.10765