IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Wenxu, Nie, Kaixuan, Du, Hang, Yin, Dong, Huang, Wei, Guo, Siqiang, Zhang, Xiaobo, Hu, Pengbo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918160148463616
author Zhou, Wenxu
Nie, Kaixuan
Du, Hang
Yin, Dong
Huang, Wei
Guo, Siqiang
Zhang, Xiaobo
Hu, Pengbo
author_facet Zhou, Wenxu
Nie, Kaixuan
Du, Hang
Yin, Dong
Huang, Wei
Guo, Siqiang
Zhang, Xiaobo
Hu, Pengbo
contents In this study, we present IL3D, a large-scale dataset meticulously designed for large language model (LLM)-driven 3D scene generation, addressing the pressing demand for diverse, high-quality training data in indoor layout design. Comprising 27,816 indoor layouts across 18 prevalent room types and a library of 29,215 high-fidelity 3D object assets, IL3D is enriched with instance-level natural language annotations to support robust multimodal learning for vision-language tasks. We establish rigorous benchmarks to evaluate LLM-driven scene generation. Experimental results show that supervised fine-tuning (SFT) of LLMs on IL3D significantly improves generalization and surpasses the performance of SFT on other datasets. IL3D offers flexible multimodal data export capabilities, including point clouds, 3D bounding boxes, multiview images, depth maps, normal maps, and semantic masks, enabling seamless adaptation to various visual tasks. As a versatile and robust resource, IL3D significantly advances research in 3D scene generation and embodied intelligence, by providing high-fidelity scene data to support environment perception tasks of embodied agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation
Zhou, Wenxu
Nie, Kaixuan
Du, Hang
Yin, Dong
Huang, Wei
Guo, Siqiang
Zhang, Xiaobo
Hu, Pengbo
Computer Vision and Pattern Recognition
In this study, we present IL3D, a large-scale dataset meticulously designed for large language model (LLM)-driven 3D scene generation, addressing the pressing demand for diverse, high-quality training data in indoor layout design. Comprising 27,816 indoor layouts across 18 prevalent room types and a library of 29,215 high-fidelity 3D object assets, IL3D is enriched with instance-level natural language annotations to support robust multimodal learning for vision-language tasks. We establish rigorous benchmarks to evaluate LLM-driven scene generation. Experimental results show that supervised fine-tuning (SFT) of LLMs on IL3D significantly improves generalization and surpasses the performance of SFT on other datasets. IL3D offers flexible multimodal data export capabilities, including point clouds, 3D bounding boxes, multiview images, depth maps, normal maps, and semantic masks, enabling seamless adaptation to various visual tasks. As a versatile and robust resource, IL3D significantly advances research in 3D scene generation and embodied intelligence, by providing high-fidelity scene data to support environment perception tasks of embodied agents.
title IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.12095