ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Jun-Kun, Bulò, Samuel Rota, Müller, Norman, Porzi, Lorenzo, Kontschieder, Peter, Wang, Yu-Xiong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909223408893952
author Chen, Jun-Kun
Bulò, Samuel Rota
Müller, Norman
Porzi, Lorenzo
Kontschieder, Peter
Wang, Yu-Xiong
author_facet Chen, Jun-Kun
Bulò, Samuel Rota
Müller, Norman
Porzi, Lorenzo
Kontschieder, Peter
Wang, Yu-Xiong
contents This paper proposes ConsistDreamer - a novel framework that lifts 2D diffusion models with 3D awareness and 3D consistency, thus enabling high-fidelity instruction-guided scene editing. To overcome the fundamental limitation of missing 3D consistency in 2D diffusion models, our key insight is to introduce three synergetic strategies that augment the input of the 2D diffusion model to become 3D-aware and to explicitly enforce 3D consistency during the training process. Specifically, we design surrounding views as context-rich input for the 2D diffusion model, and generate 3D-consistent, structured noise instead of image-independent noise. Moreover, we introduce self-supervised consistency-enforcing training within the per-scene editing procedure. Extensive evaluation shows that our ConsistDreamer achieves state-of-the-art performance for instruction-guided scene editing across various scenes and editing instructions, particularly in complicated large-scale indoor scenes from ScanNet++, with significantly improved sharpness and fine-grained textures. Notably, ConsistDreamer stands as the first work capable of successfully editing complex (e.g., plaid/checkered) patterns. Our project page is at immortalco.github.io/ConsistDreamer.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09404
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene Editing
Chen, Jun-Kun
Bulò, Samuel Rota
Müller, Norman
Porzi, Lorenzo
Kontschieder, Peter
Wang, Yu-Xiong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
This paper proposes ConsistDreamer - a novel framework that lifts 2D diffusion models with 3D awareness and 3D consistency, thus enabling high-fidelity instruction-guided scene editing. To overcome the fundamental limitation of missing 3D consistency in 2D diffusion models, our key insight is to introduce three synergetic strategies that augment the input of the 2D diffusion model to become 3D-aware and to explicitly enforce 3D consistency during the training process. Specifically, we design surrounding views as context-rich input for the 2D diffusion model, and generate 3D-consistent, structured noise instead of image-independent noise. Moreover, we introduce self-supervised consistency-enforcing training within the per-scene editing procedure. Extensive evaluation shows that our ConsistDreamer achieves state-of-the-art performance for instruction-guided scene editing across various scenes and editing instructions, particularly in complicated large-scale indoor scenes from ScanNet++, with significantly improved sharpness and fine-grained textures. Notably, ConsistDreamer stands as the first work capable of successfully editing complex (e.g., plaid/checkered) patterns. Our project page is at immortalco.github.io/ConsistDreamer.
title ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.09404