SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Songchun, Xu, Huiyao, Guo, Sitong, Xie, Zhongwei, Bao, Hujun, Xu, Weiwei, Zou, Changqing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911051231002624
author Zhang, Songchun
Xu, Huiyao
Guo, Sitong
Xie, Zhongwei
Bao, Hujun
Xu, Weiwei
Zou, Changqing
author_facet Zhang, Songchun
Xu, Huiyao
Guo, Sitong
Xie, Zhongwei
Bao, Hujun
Xu, Weiwei
Zou, Changqing
contents Novel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of reconstructing photorealistic 3D scenes from sparse or single-view inputs. We introduce SpatialCrafter, a framework that leverages the rich knowledge in video diffusion models to generate plausible additional observations, thereby alleviating reconstruction ambiguity. Through a trainable camera encoder and an epipolar attention mechanism for explicit geometric constraints, we achieve precise camera control and 3D consistency, further reinforced by a unified scale estimation strategy to handle scale discrepancies across datasets. Furthermore, by integrating monocular depth priors with semantic features in the video latent space, our framework directly regresses 3D Gaussian primitives and efficiently processes long-sequence features using a hybrid network structure. Extensive experiments show our method enhances sparse view reconstruction and restores the realistic appearance of 3D scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11992
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
Zhang, Songchun
Xu, Huiyao
Guo, Sitong
Xie, Zhongwei
Bao, Hujun
Xu, Weiwei
Zou, Changqing
Computer Vision and Pattern Recognition
Novel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of reconstructing photorealistic 3D scenes from sparse or single-view inputs. We introduce SpatialCrafter, a framework that leverages the rich knowledge in video diffusion models to generate plausible additional observations, thereby alleviating reconstruction ambiguity. Through a trainable camera encoder and an epipolar attention mechanism for explicit geometric constraints, we achieve precise camera control and 3D consistency, further reinforced by a unified scale estimation strategy to handle scale discrepancies across datasets. Furthermore, by integrating monocular depth priors with semantic features in the video latent space, our framework directly regresses 3D Gaussian primitives and efficiently processes long-sequence features using a hybrid network structure. Extensive experiments show our method enhances sparse view reconstruction and restores the realistic appearance of 3D scenes.
title SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.11992