Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guizilini, Vitor, Irshad, Muhammad Zubair, Chen, Dian, Shakhnarovich, Greg, Ambrus, Rares
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916591810117632
author Guizilini, Vitor
Irshad, Muhammad Zubair
Chen, Dian
Shakhnarovich, Greg
Ambrus, Rares
author_facet Guizilini, Vitor
Irshad, Muhammad Zubair
Chen, Dian
Shakhnarovich, Greg
Ambrus, Rares
contents Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of direct pixel-level generation of images and depth maps from novel viewpoints, given an arbitrary number of input views. Our method uses raymap conditioning to both augment visual features with spatial information from different viewpoints, as well as to guide the generation of images and depth maps from novel views. A key aspect of our approach is the multi-task generation of images and depth maps, using learnable task embeddings to guide the diffusion process towards specific modalities. We train this model on a collection of more than 60 million multi-view samples from publicly available datasets, and propose techniques to enable efficient and consistent learning in such diverse conditions. We also propose a novel strategy that enables the efficient training of larger models by incrementally fine-tuning smaller ones, with promising scaling behavior. Through extensive experiments, we report state-of-the-art results in multiple novel view synthesis benchmarks, as well as multi-view stereo and video depth estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18804
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion
Guizilini, Vitor
Irshad, Muhammad Zubair
Chen, Dian
Shakhnarovich, Greg
Ambrus, Rares
Computer Vision and Pattern Recognition
Machine Learning
Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of direct pixel-level generation of images and depth maps from novel viewpoints, given an arbitrary number of input views. Our method uses raymap conditioning to both augment visual features with spatial information from different viewpoints, as well as to guide the generation of images and depth maps from novel views. A key aspect of our approach is the multi-task generation of images and depth maps, using learnable task embeddings to guide the diffusion process towards specific modalities. We train this model on a collection of more than 60 million multi-view samples from publicly available datasets, and propose techniques to enable efficient and consistent learning in such diverse conditions. We also propose a novel strategy that enables the efficient training of larger models by incrementally fine-tuning smaller ones, with promising scaling behavior. Through extensive experiments, we report state-of-the-art results in multiple novel view synthesis benchmarks, as well as multi-view stereo and video depth estimation.
title Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2501.18804