VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Jimin, Zhang, Wenyuan, Zhou, Junsheng, Huang, Zian, Shi, Kanle, Xu, Shenkun, Liu, Yu-Shen, Han, Zhizhong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909035159093248
author Tang, Jimin
Zhang, Wenyuan
Zhou, Junsheng
Huang, Zian
Shi, Kanle
Xu, Shenkun
Liu, Yu-Shen
Han, Zhizhong
author_facet Tang, Jimin
Zhang, Wenyuan
Zhou, Junsheng
Huang, Zian
Shi, Kanle
Xu, Shenkun
Liu, Yu-Shen
Han, Zhizhong
contents Gaussian Splatting has achieved remarkable progress in multi-view surface reconstruction, yet it exhibits notable degradation when only few views are available. Although recent efforts alleviate this issue by enhancing multi-view consistency to produce plausible surfaces, they struggle to infer unseen, occluded, or weakly constrained regions beyond the input coverage. To address this limitation, we present VidSplat, a training-free generative reconstruction framework that leverages powerful video diffusion priors to iteratively synthesize novel views that compensate for missing input coverage, and thereby recover complete 3D scenes from sparse inputs. Specifically, we tackle two key challenges that enable the effective integration of generation and reconstruction. First, for 3D consistent generation, we elaborate a training-free, stage-wise denoising strategy that adaptively guides the denoising direction toward the underlying geometry using the rendered RGB and mask images. Second, to enhance the reconstruction, we develop an iterative mechanism that samples camera trajectories, explores unobserved regions, synthesizes novel views, and supplements training through confidence weighted refinement. VidSplat performs robustly to sparse input and even a single image. Extensive experiments on widely used benchmarks demonstrate our superior performance in sparse-view scene reconstruction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11424
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors
Tang, Jimin
Zhang, Wenyuan
Zhou, Junsheng
Huang, Zian
Shi, Kanle
Xu, Shenkun
Liu, Yu-Shen
Han, Zhizhong
Computer Vision and Pattern Recognition
Gaussian Splatting has achieved remarkable progress in multi-view surface reconstruction, yet it exhibits notable degradation when only few views are available. Although recent efforts alleviate this issue by enhancing multi-view consistency to produce plausible surfaces, they struggle to infer unseen, occluded, or weakly constrained regions beyond the input coverage. To address this limitation, we present VidSplat, a training-free generative reconstruction framework that leverages powerful video diffusion priors to iteratively synthesize novel views that compensate for missing input coverage, and thereby recover complete 3D scenes from sparse inputs. Specifically, we tackle two key challenges that enable the effective integration of generation and reconstruction. First, for 3D consistent generation, we elaborate a training-free, stage-wise denoising strategy that adaptively guides the denoising direction toward the underlying geometry using the rendered RGB and mask images. Second, to enhance the reconstruction, we develop an iterative mechanism that samples camera trajectories, explores unobserved regions, synthesizes novel views, and supplements training through confidence weighted refinement. VidSplat performs robustly to sparse input and even a single image. Extensive experiments on widely used benchmarks demonstrate our superior performance in sparse-view scene reconstruction.
title VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.11424