Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Xiaoming, Huang, Xu, Xie, Qinghongbing, Deng, Zhi, Yu, Junsheng, Guan, Yirui, Liu, Zhongyuan, Zhu, Lin, Zhao, Qijun, Liu, Ligang, Zeng, Long
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908599535534080
author Zhu, Xiaoming
Huang, Xu
Xie, Qinghongbing
Deng, Zhi
Yu, Junsheng
Guan, Yirui
Liu, Zhongyuan
Zhu, Lin
Zhao, Qijun
Liu, Ligang
Zeng, Long
author_facet Zhu, Xiaoming
Huang, Xu
Xie, Qinghongbing
Deng, Zhi
Yu, Junsheng
Guan, Yirui
Liu, Zhongyuan
Zhu, Lin
Zhao, Qijun
Liu, Ligang
Zeng, Long
contents Generating artistic and coherent 3D scene layouts is crucial in digital content creation. Traditional optimization-based methods are often constrained by cumbersome manual rules, while deep generative models face challenges in producing content with richness and diversity. Furthermore, approaches that utilize large language models frequently lack robustness and fail to accurately capture complex spatial relationships. To address these challenges, this paper presents a novel vision-guided 3D layout generation system. We first construct a high-quality asset library containing 2,037 scene assets and 147 3D scene layouts. Subsequently, we employ an image generation model to expand prompt representations into images, fine-tuning it to align with our asset library. We then develop a robust image parsing module to recover the 3D layout of scenes based on visual semantics and geometric information. Finally, we optimize the scene layout using scene graphs and overall visual semantics to ensure logical coherence and alignment with the images. Extensive user testing demonstrates that our algorithm significantly outperforms existing methods in terms of layout richness and quality. The code and dataset will be available at https://github.com/HiHiAllen/Imaginarium.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15564
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation
Zhu, Xiaoming
Huang, Xu
Xie, Qinghongbing
Deng, Zhi
Yu, Junsheng
Guan, Yirui
Liu, Zhongyuan
Zhu, Lin
Zhao, Qijun
Liu, Ligang
Zeng, Long
Computer Vision and Pattern Recognition
Generating artistic and coherent 3D scene layouts is crucial in digital content creation. Traditional optimization-based methods are often constrained by cumbersome manual rules, while deep generative models face challenges in producing content with richness and diversity. Furthermore, approaches that utilize large language models frequently lack robustness and fail to accurately capture complex spatial relationships. To address these challenges, this paper presents a novel vision-guided 3D layout generation system. We first construct a high-quality asset library containing 2,037 scene assets and 147 3D scene layouts. Subsequently, we employ an image generation model to expand prompt representations into images, fine-tuning it to align with our asset library. We then develop a robust image parsing module to recover the 3D layout of scenes based on visual semantics and geometric information. Finally, we optimize the scene layout using scene graphs and overall visual semantics to ensure logical coherence and alignment with the images. Extensive user testing demonstrates that our algorithm significantly outperforms existing methods in terms of layout richness and quality. The code and dataset will be available at https://github.com/HiHiAllen/Imaginarium.
title Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.15564