Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wulff, Philipp, Wimbauer, Felix, Muhle, Dominik, Cremers, Daniel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918113649360896
author Wulff, Philipp
Wimbauer, Felix
Muhle, Dominik
Cremers, Daniel
author_facet Wulff, Philipp
Wimbauer, Felix
Muhle, Dominik
Cremers, Daniel
contents Volumetric scene reconstruction from a single image is crucial for a broad range of applications like autonomous driving and robotics. Recent volumetric reconstruction methods achieve impressive results, but generally require expensive 3D ground truth or multi-view supervision. We propose to leverage pre-trained 2D diffusion models and depth prediction models to generate synthetic scene geometry from a single image. This can then be used to distill a feed-forward scene reconstruction model. Our experiments on the challenging KITTI-360 and Waymo datasets demonstrate that our method matches or outperforms state-of-the-art baselines that use multi-view supervision, and offers unique advantages, for example regarding dynamic scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02323
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
Wulff, Philipp
Wimbauer, Felix
Muhle, Dominik
Cremers, Daniel
Computer Vision and Pattern Recognition
Volumetric scene reconstruction from a single image is crucial for a broad range of applications like autonomous driving and robotics. Recent volumetric reconstruction methods achieve impressive results, but generally require expensive 3D ground truth or multi-view supervision. We propose to leverage pre-trained 2D diffusion models and depth prediction models to generate synthetic scene geometry from a single image. This can then be used to distill a feed-forward scene reconstruction model. Our experiments on the challenging KITTI-360 and Waymo datasets demonstrate that our method matches or outperforms state-of-the-art baselines that use multi-view supervision, and offers unique advantages, for example regarding dynamic scenes.
title Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.02323