MeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Xuyang, Zhai, Zhijun, Zhou, Kaixuan, Wang, Zengmao, He, Jianan, Wang, Dong, Zhang, Yanfeng, Sun, mingwei, Westermann, Rüdiger, Schindler, Konrad, Meng, Liqiu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915705868255232
author Chen, Xuyang
Zhai, Zhijun
Zhou, Kaixuan
Wang, Zengmao
He, Jianan
Wang, Dong
Zhang, Yanfeng
Sun, mingwei
Westermann, Rüdiger
Schindler, Konrad
Meng, Liqiu
author_facet Chen, Xuyang
Zhai, Zhijun
Zhou, Kaixuan
Wang, Zengmao
He, Jianan
Wang, Dong
Zhang, Yanfeng
Sun, mingwei
Westermann, Rüdiger
Schindler, Konrad
Meng, Liqiu
contents Mesh models have become increasingly accessible for numerous cities; however, the lack of realistic textures restricts their application in virtual urban navigation and autonomous driving. To address this, this paper proposes MeSS (Meshbased Scene Synthesis) for generating high-quality, styleconsistent outdoor scenes with city mesh models serving as the geometric prior. While image and video diffusion models can leverage spatial layouts (such as depth maps or HD maps) as control conditions to generate street-level perspective views, they are not directly applicable to 3D scene generation. Video diffusion models excel at synthesizing consistent view sequences that depict scenes but often struggle to adhere to predefined camera paths or align accurately with rendered control videos. In contrast, image diffusion models, though unable to guarantee cross-view visual consistency, can produce more geometry-aligned results when combined with ControlNet. Building on this insight, our approach enhances image diffusion models by improving cross-view consistency. The pipeline comprises three key stages: first, we generate geometrically consistent sparse views using Cascaded Outpainting ControlNets; second, we propagate denser intermediate views via a component dubbed AGInpaint; and third, we globally eliminate visual inconsistencies (e.g., varying exposure) using the GCAlign module. Concurrently with generation, a 3D Gaussian Splatting (3DGS) scene is reconstructed by initializing Gaussian balls on the mesh surface. Our method outperforms existing approaches in both geometric alignment and generation quality. Once synthesized, the scene can be rendered in diverse styles through relighting and style transfer techniques. project page: https://albertchen98.github.io/mess/
format Preprint
id arxiv_https___arxiv_org_abs_2508_15169
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusion
Chen, Xuyang
Zhai, Zhijun
Zhou, Kaixuan
Wang, Zengmao
He, Jianan
Wang, Dong
Zhang, Yanfeng
Sun, mingwei
Westermann, Rüdiger
Schindler, Konrad
Meng, Liqiu
Computer Vision and Pattern Recognition
Mesh models have become increasingly accessible for numerous cities; however, the lack of realistic textures restricts their application in virtual urban navigation and autonomous driving. To address this, this paper proposes MeSS (Meshbased Scene Synthesis) for generating high-quality, styleconsistent outdoor scenes with city mesh models serving as the geometric prior. While image and video diffusion models can leverage spatial layouts (such as depth maps or HD maps) as control conditions to generate street-level perspective views, they are not directly applicable to 3D scene generation. Video diffusion models excel at synthesizing consistent view sequences that depict scenes but often struggle to adhere to predefined camera paths or align accurately with rendered control videos. In contrast, image diffusion models, though unable to guarantee cross-view visual consistency, can produce more geometry-aligned results when combined with ControlNet. Building on this insight, our approach enhances image diffusion models by improving cross-view consistency. The pipeline comprises three key stages: first, we generate geometrically consistent sparse views using Cascaded Outpainting ControlNets; second, we propagate denser intermediate views via a component dubbed AGInpaint; and third, we globally eliminate visual inconsistencies (e.g., varying exposure) using the GCAlign module. Concurrently with generation, a 3D Gaussian Splatting (3DGS) scene is reconstructed by initializing Gaussian balls on the mesh surface. Our method outperforms existing approaches in both geometric alignment and generation quality. Once synthesized, the scene can be rendered in diverse styles through relighting and style transfer techniques. project page: https://albertchen98.github.io/mess/
title MeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.15169