Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Qirui, Iliash, Denys, Ritchie, Daniel, Savva, Manolis, Chang, Angel X.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912275995033600
author Wu, Qirui
Iliash, Denys
Ritchie, Daniel
Savva, Manolis
Chang, Angel X.
author_facet Wu, Qirui
Iliash, Denys
Ritchie, Daniel
Savva, Manolis
Chang, Angel X.
contents Reconstructing structured 3D scenes from RGB images using CAD objects unlocks efficient and compact scene representations that maintain compositionality and interactability. Existing works propose training-heavy methods relying on either expensive yet inaccurate real-world annotations or controllable yet monotonous synthetic data that do not generalize well to unseen objects or domains. We present Diorama, the first zero-shot open-world system that holistically models 3D scenes from single-view RGB observations without requiring end-to-end training or human annotations. We show the feasibility of our approach by decomposing the problem into subtasks and introduce robust, generalizable solutions to each: architecture reconstruction, 3D shape retrieval, object pose estimation, and scene layout optimization. We evaluate our system on both synthetic and real-world data to show we significantly outperform baselines from prior work. We also demonstrate generalization to internet images and the text-to-scene task.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19492
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene Modeling
Wu, Qirui
Iliash, Denys
Ritchie, Daniel
Savva, Manolis
Chang, Angel X.
Computer Vision and Pattern Recognition
Machine Learning
Reconstructing structured 3D scenes from RGB images using CAD objects unlocks efficient and compact scene representations that maintain compositionality and interactability. Existing works propose training-heavy methods relying on either expensive yet inaccurate real-world annotations or controllable yet monotonous synthetic data that do not generalize well to unseen objects or domains. We present Diorama, the first zero-shot open-world system that holistically models 3D scenes from single-view RGB observations without requiring end-to-end training or human annotations. We show the feasibility of our approach by decomposing the problem into subtasks and introduce robust, generalizable solutions to each: architecture reconstruction, 3D shape retrieval, object pose estimation, and scene layout optimization. We evaluate our system on both synthetic and real-world data to show we significantly outperform baselines from prior work. We also demonstrate generalization to internet images and the text-to-scene task.
title Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene Modeling
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.19492