Map2World: Segment Map Conditioned Text to 3D World Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chung, Jaeyoung, Lee, Suyoung, Xiang, Jianfeng, Yang, Jiaolong, Lee, Kyoung Mu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915972769644544
author Chung, Jaeyoung
Lee, Suyoung
Xiang, Jianfeng
Yang, Jiaolong
Lee, Kyoung Mu
author_facet Chung, Jaeyoung
Lee, Suyoung
Xiang, Jianfeng
Yang, Jiaolong
Lee, Kyoung Mu
contents 3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D world generation conditioned on user-defined segment maps of arbitrary shapes and scales, ensuring global-scale consistency and flexibility across expansive environments. To further enhance the quality, we propose a detail enhancer network that generates fine details of the world. The detail enhancer enables the addition of fine-grained details without compromising overall scene coherence by incorporating global structure information. We design the entire pipeline to leverage strong priors from asset generators, achieving robust generalization across diverse domains, even under limited training data for scene generation. Extensive experiments demonstrate that our method significantly outperforms existing approaches in user-controllability, scale consistency, and content coherence, enabling users to generate 3D worlds under more complex conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00781
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Map2World: Segment Map Conditioned Text to 3D World Generation
Chung, Jaeyoung
Lee, Suyoung
Xiang, Jianfeng
Yang, Jiaolong
Lee, Kyoung Mu
Computer Vision and Pattern Recognition
3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D world generation conditioned on user-defined segment maps of arbitrary shapes and scales, ensuring global-scale consistency and flexibility across expansive environments. To further enhance the quality, we propose a detail enhancer network that generates fine details of the world. The detail enhancer enables the addition of fine-grained details without compromising overall scene coherence by incorporating global structure information. We design the entire pipeline to leverage strong priors from asset generators, achieving robust generalization across diverse domains, even under limited training data for scene generation. Extensive experiments demonstrate that our method significantly outperforms existing approaches in user-controllability, scale consistency, and content coherence, enabling users to generate 3D worlds under more complex conditions.
title Map2World: Segment Map Conditioned Text to 3D World Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.00781