Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yixuan, Luo, Zhen, Gan, Wanshui, Hao, Jinkun, Lu, Junru, Yan, Jinghao, Lyu, Zhaoyang, Xu, Xudong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918509313785856
author Yang, Yixuan
Luo, Zhen
Gan, Wanshui
Hao, Jinkun
Lu, Junru
Yan, Jinghao
Lyu, Zhaoyang
Xu, Xudong
author_facet Yang, Yixuan
Luo, Zhen
Gan, Wanshui
Hao, Jinkun
Lu, Junru
Yan, Jinghao
Lyu, Zhaoyang
Xu, Xudong
contents Designing realistic and functional 3D indoor rooms is essential for a wide range of applications, including interior design, virtual reality, gaming, and embodied AI. While recent MLLM-based approaches have shown great potential for 3D room synthesis from textual descriptions or reference images, text-based methods struggle to capture precise spatial information, and existing image-conditioned agents suffer from instability and infinite looping when tasked with holistic room generation from top-down views. To address these limitations, we propose Code-as-Room, an MLLM-based agentic framework equipped with a structured execution harness, which represents 3D rooms with Blender codes. Given a top-down room image, the framework parses the reference image to extract scene elements and their spatial relationships, and synthesizes executable Blender code for geometry, materials, and lighting in a principled, multi-stage pipeline. A cross-stage memory module is maintained throughout to mitigate context forgetting inherent to existing agent-based frameworks. We further introduce a dedicated benchmark for code-based 3D room synthesis, encompassing various evaluation protocols. Based on our benchmark, comprehensive comparisons against existing agent-based methods are conducted to validate the effectiveness of our proposed execution harness.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18451
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis
Yang, Yixuan
Luo, Zhen
Gan, Wanshui
Hao, Jinkun
Lu, Junru
Yan, Jinghao
Lyu, Zhaoyang
Xu, Xudong
Computer Vision and Pattern Recognition
Graphics
Designing realistic and functional 3D indoor rooms is essential for a wide range of applications, including interior design, virtual reality, gaming, and embodied AI. While recent MLLM-based approaches have shown great potential for 3D room synthesis from textual descriptions or reference images, text-based methods struggle to capture precise spatial information, and existing image-conditioned agents suffer from instability and infinite looping when tasked with holistic room generation from top-down views. To address these limitations, we propose Code-as-Room, an MLLM-based agentic framework equipped with a structured execution harness, which represents 3D rooms with Blender codes. Given a top-down room image, the framework parses the reference image to extract scene elements and their spatial relationships, and synthesizes executable Blender code for geometry, materials, and lighting in a principled, multi-stage pipeline. A cross-stage memory module is maintained throughout to mitigate context forgetting inherent to existing agent-based frameworks. We further introduce a dedicated benchmark for code-based 3D room synthesis, encompassing various evaluation protocols. Based on our benchmark, comprehensive comparisons against existing agent-based methods are conducted to validate the effectiveness of our proposed execution harness.
title Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2605.18451