MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zihan, Tan, Jeff, Khurana, Tarasha, Peri, Neehar, Ramanan, Deva
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914360190828544
author Wang, Zihan
Tan, Jeff
Khurana, Tarasha
Peri, Neehar
Ramanan, Deva
author_facet Wang, Zihan
Tan, Jeff
Khurana, Tarasha
Peri, Neehar
Ramanan, Deva
contents We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panoptic Studio). Such multi-view setups are prohibitively expensive to build and cannot capture diverse scenes in-the-wild. In contrast, we aim to reconstruct dynamic human behaviors, such as repairing a bike or dancing, from a small set of sparse-view cameras with complete scene coverage (e.g. four equidistant inward-facing static cameras). We find that dense multi-view reconstruction methods struggle to adapt to this sparse-view setup due to limited overlap between viewpoints. To address these limitations, we carefully align independent monocular reconstructions of each camera to produce time- and view-consistent dynamic scene reconstructions. Extensive experiments on PanopticStudio and Ego-Exo4D demonstrate that our method achieves higher quality reconstructions than prior art, particularly when rendering novel views. Code, data, and data-processing scripts are available on https://github.com/Z1hanW/MonoFusion.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23782
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
Wang, Zihan
Tan, Jeff
Khurana, Tarasha
Peri, Neehar
Ramanan, Deva
Computer Vision and Pattern Recognition
We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panoptic Studio). Such multi-view setups are prohibitively expensive to build and cannot capture diverse scenes in-the-wild. In contrast, we aim to reconstruct dynamic human behaviors, such as repairing a bike or dancing, from a small set of sparse-view cameras with complete scene coverage (e.g. four equidistant inward-facing static cameras). We find that dense multi-view reconstruction methods struggle to adapt to this sparse-view setup due to limited overlap between viewpoints. To address these limitations, we carefully align independent monocular reconstructions of each camera to produce time- and view-consistent dynamic scene reconstructions. Extensive experiments on PanopticStudio and Ego-Exo4D demonstrate that our method achieves higher quality reconstructions than prior art, particularly when rendering novel views. Code, data, and data-processing scripts are available on https://github.com/Z1hanW/MonoFusion.
title MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.23782