FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhifei, Zhai, Guangyao, Lu, Keyang, Yin, YuYang, Zhang, Chao, Xiao, Zhen, Long, Jieyi, Navab, Nassir, Wang, Yikai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915876464230400
author Yang, Zhifei
Zhai, Guangyao
Lu, Keyang
Yin, YuYang
Zhang, Chao
Xiao, Zhen
Long, Jieyi
Navab, Nassir
Wang, Yikai
author_facet Yang, Zhifei
Zhai, Guangyao
Lu, Keyang
Yin, YuYang
Zhang, Chao
Xiao, Zhen
Long, Jieyi
Navab, Nassir
Wang, Yikai
contents Scene generation has extensive industrial applications, demanding both high realism and precise control over geometry and appearance. Language-driven retrieval methods compose plausible scenes from a large object database, but overlook object-level control and often fail to enforce scene-level style coherence. Graph-based formulations offer higher controllability over objects and inform holistic consistency by explicitly modeling relations, yet existing methods struggle to produce high-fidelity textured results, thereby limiting their practical utility. We present FlowScene, a tri-branch scene generative model conditioned on multimodal graphs that collaboratively generates scene layouts, object shapes, and object textures. At its core lies a tight-coupled rectified flow model that exchanges object information during generation, enabling collaborative reasoning across the graph. This enables fine-grained control of objects' shapes, textures, and relations while enforcing scene-level style coherence across structure and appearance. Extensive experiments show that FlowScene outperforms both language-conditioned and graph-conditioned baselines in terms of generation realism, style consistency, and alignment with human preferences.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19598
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow
Yang, Zhifei
Zhai, Guangyao
Lu, Keyang
Yin, YuYang
Zhang, Chao
Xiao, Zhen
Long, Jieyi
Navab, Nassir
Wang, Yikai
Computer Vision and Pattern Recognition
Scene generation has extensive industrial applications, demanding both high realism and precise control over geometry and appearance. Language-driven retrieval methods compose plausible scenes from a large object database, but overlook object-level control and often fail to enforce scene-level style coherence. Graph-based formulations offer higher controllability over objects and inform holistic consistency by explicitly modeling relations, yet existing methods struggle to produce high-fidelity textured results, thereby limiting their practical utility. We present FlowScene, a tri-branch scene generative model conditioned on multimodal graphs that collaboratively generates scene layouts, object shapes, and object textures. At its core lies a tight-coupled rectified flow model that exchanges object information during generation, enabling collaborative reasoning across the graph. This enables fine-grained control of objects' shapes, textures, and relations while enforcing scene-level style coherence across structure and appearance. Extensive experiments show that FlowScene outperforms both language-conditioned and graph-conditioned baselines in terms of generation realism, style consistency, and alignment with human preferences.
title FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.19598