Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Dehui, Xu, Congsheng, Wei, Rong, Shi, Yue, Chen, Shoufa, Luo, Dingxiang, Yang, Tianshuo, Yang, Xiaokang, Sui, Wei, Qin, Yusen, Tang, Rui, Mu, Yao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913028846387200
author Wang, Dehui
Xu, Congsheng
Wei, Rong
Shi, Yue
Chen, Shoufa
Luo, Dingxiang
Yang, Tianshuo
Yang, Xiaokang
Sui, Wei
Qin, Yusen
Tang, Rui
Mu, Yao
author_facet Wang, Dehui
Xu, Congsheng
Wei, Rong
Shi, Yue
Chen, Shoufa
Luo, Dingxiang
Yang, Tianshuo
Yang, Xiaokang
Sui, Wei
Qin, Yusen
Tang, Rui
Mu, Yao
contents The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. However, existing approaches struggle to infer massive amounts of missing geometry in large unseen areas while maintaining global consistency, often producing locally plausible but globally inconsistent reconstructions. We present Rein3D, a framework that reconstructs full 360-degree indoor environments by coupling explicit 3D Gaussian Splatting (3DGS) with temporally coherent priors from video diffusion models. Our approach follows a "restore-and-refine" paradigm: we employ a radial exploration strategy to render imperfect panoramic videos along trajectories starting from the origin, effectively uncovering occluded regions from a coarse 3DGS initialization. These sequences are restored by a panoramic video-to-video diffusion model and further enhanced via video super-resolution to synthesize high-fidelity geometry and textures. Finally, these refined videos serve as pseudo-ground truths to update the global 3D Gaussian field. To support this task, we construct PanoV2V-15K, a dataset of over 15K paired clean and degraded panoramic videos for diffusion-based scene restoration. Experiments demonstrate that Rein3D produces photorealistic and globally consistent 3D scenes and significantly improves long-range camera exploration compared with existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10578
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models
Wang, Dehui
Xu, Congsheng
Wei, Rong
Shi, Yue
Chen, Shoufa
Luo, Dingxiang
Yang, Tianshuo
Yang, Xiaokang
Sui, Wei
Qin, Yusen
Tang, Rui
Mu, Yao
Computer Vision and Pattern Recognition
The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. However, existing approaches struggle to infer massive amounts of missing geometry in large unseen areas while maintaining global consistency, often producing locally plausible but globally inconsistent reconstructions. We present Rein3D, a framework that reconstructs full 360-degree indoor environments by coupling explicit 3D Gaussian Splatting (3DGS) with temporally coherent priors from video diffusion models. Our approach follows a "restore-and-refine" paradigm: we employ a radial exploration strategy to render imperfect panoramic videos along trajectories starting from the origin, effectively uncovering occluded regions from a coarse 3DGS initialization. These sequences are restored by a panoramic video-to-video diffusion model and further enhanced via video super-resolution to synthesize high-fidelity geometry and textures. Finally, these refined videos serve as pseudo-ground truths to update the global 3D Gaussian field. To support this task, we construct PanoV2V-15K, a dataset of over 15K paired clean and degraded panoramic videos for diffusion-based scene restoration. Experiments demonstrate that Rein3D produces photorealistic and globally consistent 3D scenes and significantly improves long-range camera exploration compared with existing baselines.
title Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.10578