GenFusion: Closing the Loop between Reconstruction and Generation via Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Sibo, Xu, Congrong, Huang, Binbin, Geiger, Andreas, Chen, Anpei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917971211845632
author Wu, Sibo
Xu, Congrong
Huang, Binbin
Geiger, Andreas
Chen, Anpei
author_facet Wu, Sibo
Xu, Congrong
Huang, Binbin
Geiger, Andreas
Chen, Anpei
contents Recently, 3D reconstruction and generation have demonstrated impressive novel view synthesis results, achieving high fidelity and efficiency. However, a notable conditioning gap can be observed between these two fields, e.g., scalable 3D scene reconstruction often requires densely captured views, whereas 3D generation typically relies on a single or no input view, which significantly limits their applications. We found that the source of this phenomenon lies in the misalignment between 3D constraints and generative priors. To address this problem, we propose a reconstruction-driven video diffusion model that learns to condition video frames on artifact-prone RGB-D renderings. Moreover, we propose a cyclical fusion pipeline that iteratively adds restoration frames from the generative model to the training set, enabling progressive expansion and addressing the viewpoint saturation limitations seen in previous reconstruction and generation pipelines. Our evaluation, including view synthesis from sparse view and masked input, validates the effectiveness of our approach. More details at https://genfusion.sibowu.com.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21219
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GenFusion: Closing the Loop between Reconstruction and Generation via Videos
Wu, Sibo
Xu, Congrong
Huang, Binbin
Geiger, Andreas
Chen, Anpei
Computer Vision and Pattern Recognition
Artificial Intelligence
Recently, 3D reconstruction and generation have demonstrated impressive novel view synthesis results, achieving high fidelity and efficiency. However, a notable conditioning gap can be observed between these two fields, e.g., scalable 3D scene reconstruction often requires densely captured views, whereas 3D generation typically relies on a single or no input view, which significantly limits their applications. We found that the source of this phenomenon lies in the misalignment between 3D constraints and generative priors. To address this problem, we propose a reconstruction-driven video diffusion model that learns to condition video frames on artifact-prone RGB-D renderings. Moreover, we propose a cyclical fusion pipeline that iteratively adds restoration frames from the generative model to the training set, enabling progressive expansion and addressing the viewpoint saturation limitations seen in previous reconstruction and generation pipelines. Our evaluation, including view synthesis from sparse view and masked input, validates the effectiveness of our approach. More details at https://genfusion.sibowu.com.
title GenFusion: Closing the Loop between Reconstruction and Generation via Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.21219