SceneCraft: Layout-Guided 3D Scene Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Xiuyu, Man, Yunze, Chen, Jun-Kun, Wang, Yu-Xiong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909603840655360
author Yang, Xiuyu
Man, Yunze
Chen, Jun-Kun
Wang, Yu-Xiong
author_facet Yang, Xiuyu
Man, Yunze
Chen, Jun-Kun
Wang, Yu-Xiong
contents The creation of complex 3D scenes tailored to user specifications has been a tedious and challenging task with traditional 3D modeling tools. Although some pioneering methods have achieved automatic text-to-3D generation, they are generally limited to small-scale scenes with restricted control over the shape and texture. We introduce SceneCraft, a novel method for generating detailed indoor scenes that adhere to textual descriptions and spatial layout preferences provided by users. Central to our method is a rendering-based technique, which converts 3D semantic layouts into multi-view 2D proxy maps. Furthermore, we design a semantic and depth conditioned diffusion model to generate multi-view images, which are used to learn a neural radiance field (NeRF) as the final scene representation. Without the constraints of panorama image generation, we surpass previous methods in supporting complicated indoor space generation beyond a single room, even as complicated as a whole multi-bedroom apartment with irregular shapes and layouts. Through experimental analysis, we demonstrate that our method significantly outperforms existing approaches in complex indoor scene generation with diverse textures, consistent geometry, and realistic visual quality. Code and more results are available at: https://orangesodahub.github.io/SceneCraft
format Preprint
id arxiv_https___arxiv_org_abs_2410_09049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SceneCraft: Layout-Guided 3D Scene Generation
Yang, Xiuyu
Man, Yunze
Chen, Jun-Kun
Wang, Yu-Xiong
Computer Vision and Pattern Recognition
The creation of complex 3D scenes tailored to user specifications has been a tedious and challenging task with traditional 3D modeling tools. Although some pioneering methods have achieved automatic text-to-3D generation, they are generally limited to small-scale scenes with restricted control over the shape and texture. We introduce SceneCraft, a novel method for generating detailed indoor scenes that adhere to textual descriptions and spatial layout preferences provided by users. Central to our method is a rendering-based technique, which converts 3D semantic layouts into multi-view 2D proxy maps. Furthermore, we design a semantic and depth conditioned diffusion model to generate multi-view images, which are used to learn a neural radiance field (NeRF) as the final scene representation. Without the constraints of panorama image generation, we surpass previous methods in supporting complicated indoor space generation beyond a single room, even as complicated as a whole multi-bedroom apartment with irregular shapes and layouts. Through experimental analysis, we demonstrate that our method significantly outperforms existing approaches in complex indoor scene generation with diverse textures, consistent geometry, and realistic visual quality. Code and more results are available at: https://orangesodahub.github.io/SceneCraft
title SceneCraft: Layout-Guided 3D Scene Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.09049