EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Kaizhi, Chen, Xiaotong, He, Xuehai, Gu, Jing, Li, Linjie, Yang, Zhengyuan, Lin, Kevin, Wang, Jianfeng, Wang, Lijuan, Wang, Xin Eric
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917973368766464
author Zheng, Kaizhi
Chen, Xiaotong
He, Xuehai
Gu, Jing
Li, Linjie
Yang, Zhengyuan
Lin, Kevin
Wang, Jianfeng
Wang, Lijuan
Wang, Xin Eric
author_facet Zheng, Kaizhi
Chen, Xiaotong
He, Xuehai
Gu, Jing
Li, Linjie
Yang, Zhengyuan
Lin, Kevin
Wang, Jianfeng
Wang, Lijuan
Wang, Xin Eric
contents Given the steep learning curve of professional 3D software and the time-consuming process of managing large 3D assets, language-guided 3D scene editing has significant potential in fields such as virtual reality, augmented reality, and gaming. However, recent approaches to language-guided 3D scene editing either require manual interventions or focus only on appearance modifications without supporting comprehensive scene layout changes. In response, we propose EditRoom, a unified framework capable of executing a variety of layout edits through natural language commands, without requiring manual intervention. Specifically, EditRoom leverages Large Language Models (LLMs) for command planning and generates target scenes using a diffusion-based method, enabling six types of edits: rotate, translate, scale, replace, add, and remove. To address the lack of data for language-guided 3D scene editing, we have developed an automatic pipeline to augment existing 3D scene synthesis datasets and introduced EditRoom-DB, a large-scale dataset with 83k editing pairs, for training and evaluation. Our experiments demonstrate that our approach consistently outperforms other baselines across all metrics, indicating higher accuracy and coherence in language-guided scene layout editing.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12836
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
Zheng, Kaizhi
Chen, Xiaotong
He, Xuehai
Gu, Jing
Li, Linjie
Yang, Zhengyuan
Lin, Kevin
Wang, Jianfeng
Wang, Lijuan
Wang, Xin Eric
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
Given the steep learning curve of professional 3D software and the time-consuming process of managing large 3D assets, language-guided 3D scene editing has significant potential in fields such as virtual reality, augmented reality, and gaming. However, recent approaches to language-guided 3D scene editing either require manual interventions or focus only on appearance modifications without supporting comprehensive scene layout changes. In response, we propose EditRoom, a unified framework capable of executing a variety of layout edits through natural language commands, without requiring manual intervention. Specifically, EditRoom leverages Large Language Models (LLMs) for command planning and generates target scenes using a diffusion-based method, enabling six types of edits: rotate, translate, scale, replace, add, and remove. To address the lack of data for language-guided 3D scene editing, we have developed an automatic pipeline to augment existing 3D scene synthesis datasets and introduced EditRoom-DB, a large-scale dataset with 83k editing pairs, for training and evaluation. Our experiments demonstrate that our approach consistently outperforms other baselines across all metrics, indicating higher accuracy and coherence in language-guided scene layout editing.
title EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2410.12836