Coherent 3D Scene Diffusion From a Single RGB Image

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dahnert, Manuel, Dai, Angela, Müller, Norman, Nießner, Matthias
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915063017766912
author Dahnert, Manuel
Dai, Angela
Müller, Norman
Nießner, Matthias
author_facet Dahnert, Manuel
Dai, Angela
Müller, Norman
Nießner, Matthias
contents We present a novel diffusion-based approach for coherent 3D scene reconstruction from a single RGB image. Our method utilizes an image-conditioned 3D scene diffusion model to simultaneously denoise the 3D poses and geometries of all objects within the scene. Motivated by the ill-posed nature of the task and to obtain consistent scene reconstruction results, we learn a generative scene prior by conditioning on all scene objects simultaneously to capture the scene context and by allowing the model to learn inter-object relationships throughout the diffusion process. We further propose an efficient surface alignment loss to facilitate training even in the absence of full ground-truth annotation, which is common in publicly available datasets. This loss leverages an expressive shape representation, which enables direct point sampling from intermediate shape predictions. By framing the task of single RGB image 3D scene reconstruction as a conditional diffusion process, our approach surpasses current state-of-the-art methods, achieving a 12.04% improvement in AP3D on SUN RGB-D and a 13.43% increase in F-Score on Pix3D.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10294
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Coherent 3D Scene Diffusion From a Single RGB Image
Dahnert, Manuel
Dai, Angela
Müller, Norman
Nießner, Matthias
Computer Vision and Pattern Recognition
We present a novel diffusion-based approach for coherent 3D scene reconstruction from a single RGB image. Our method utilizes an image-conditioned 3D scene diffusion model to simultaneously denoise the 3D poses and geometries of all objects within the scene. Motivated by the ill-posed nature of the task and to obtain consistent scene reconstruction results, we learn a generative scene prior by conditioning on all scene objects simultaneously to capture the scene context and by allowing the model to learn inter-object relationships throughout the diffusion process. We further propose an efficient surface alignment loss to facilitate training even in the absence of full ground-truth annotation, which is common in publicly available datasets. This loss leverages an expressive shape representation, which enables direct point sampling from intermediate shape predictions. By framing the task of single RGB image 3D scene reconstruction as a conditional diffusion process, our approach surpasses current state-of-the-art methods, achieving a 12.04% improvement in AP3D on SUN RGB-D and a 13.43% increase in F-Score on Pix3D.
title Coherent 3D Scene Diffusion From a Single RGB Image
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.10294