FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915876464230400 |
|---|---|
| author | Yang, Zhifei Zhai, Guangyao Lu, Keyang Yin, YuYang Zhang, Chao Xiao, Zhen Long, Jieyi Navab, Nassir Wang, Yikai |
| author_facet | Yang, Zhifei Zhai, Guangyao Lu, Keyang Yin, YuYang Zhang, Chao Xiao, Zhen Long, Jieyi Navab, Nassir Wang, Yikai |
| contents | Scene generation has extensive industrial applications, demanding both high realism and precise control over geometry and appearance. Language-driven retrieval methods compose plausible scenes from a large object database, but overlook object-level control and often fail to enforce scene-level style coherence. Graph-based formulations offer higher controllability over objects and inform holistic consistency by explicitly modeling relations, yet existing methods struggle to produce high-fidelity textured results, thereby limiting their practical utility. We present FlowScene, a tri-branch scene generative model conditioned on multimodal graphs that collaboratively generates scene layouts, object shapes, and object textures. At its core lies a tight-coupled rectified flow model that exchanges object information during generation, enabling collaborative reasoning across the graph. This enables fine-grained control of objects' shapes, textures, and relations while enforcing scene-level style coherence across structure and appearance. Extensive experiments show that FlowScene outperforms both language-conditioned and graph-conditioned baselines in terms of generation realism, style consistency, and alignment with human preferences. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_19598 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow Yang, Zhifei Zhai, Guangyao Lu, Keyang Yin, YuYang Zhang, Chao Xiao, Zhen Long, Jieyi Navab, Nassir Wang, Yikai Computer Vision and Pattern Recognition Scene generation has extensive industrial applications, demanding both high realism and precise control over geometry and appearance. Language-driven retrieval methods compose plausible scenes from a large object database, but overlook object-level control and often fail to enforce scene-level style coherence. Graph-based formulations offer higher controllability over objects and inform holistic consistency by explicitly modeling relations, yet existing methods struggle to produce high-fidelity textured results, thereby limiting their practical utility. We present FlowScene, a tri-branch scene generative model conditioned on multimodal graphs that collaboratively generates scene layouts, object shapes, and object textures. At its core lies a tight-coupled rectified flow model that exchanges object information during generation, enabling collaborative reasoning across the graph. This enables fine-grained control of objects' shapes, textures, and relations while enforcing scene-level style coherence across structure and appearance. Extensive experiments show that FlowScene outperforms both language-conditioned and graph-conditioned baselines in terms of generation realism, style consistency, and alignment with human preferences. |
| title | FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.19598 |