SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Xiaohao, Goel, Divyam, Chang, Angel X.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912573868212224
author Sun, Xiaohao
Goel, Divyam
Chang, Angel X.
author_facet Sun, Xiaohao
Goel, Divyam
Chang, Angel X.
contents We present SemLayoutDiff, a unified model for synthesizing diverse 3D indoor scenes across multiple room types. The model introduces a scene layout representation combining a top-down semantic map and attributes for each object. Unlike prior approaches, which cannot condition on architectural constraints, SemLayoutDiff employs a categorical diffusion model capable of conditioning scene synthesis explicitly on room masks. It first generates a coherent semantic map, followed by a cross-attention-based network to predict furniture placements that respect the synthesized layout. Our method also accounts for architectural elements such as doors and windows, ensuring that generated furniture arrangements remain practical and unobstructed. Experiments on the 3D-FRONT dataset show that SemLayoutDiff produces spatially coherent, realistic, and varied scenes, outperforming previous methods.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18597
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis
Sun, Xiaohao
Goel, Divyam
Chang, Angel X.
Graphics
Computer Vision and Pattern Recognition
We present SemLayoutDiff, a unified model for synthesizing diverse 3D indoor scenes across multiple room types. The model introduces a scene layout representation combining a top-down semantic map and attributes for each object. Unlike prior approaches, which cannot condition on architectural constraints, SemLayoutDiff employs a categorical diffusion model capable of conditioning scene synthesis explicitly on room masks. It first generates a coherent semantic map, followed by a cross-attention-based network to predict furniture placements that respect the synthesized layout. Our method also accounts for architectural elements such as doors and windows, ensuring that generated furniture arrangements remain practical and unobstructed. Experiments on the 3D-FRONT dataset show that SemLayoutDiff produces spatially coherent, realistic, and varied scenes, outperforming previous methods.
title SemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesis
topic Graphics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.18597