Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918113649360896 |
|---|---|
| author | Wulff, Philipp Wimbauer, Felix Muhle, Dominik Cremers, Daniel |
| author_facet | Wulff, Philipp Wimbauer, Felix Muhle, Dominik Cremers, Daniel |
| contents | Volumetric scene reconstruction from a single image is crucial for a broad range of applications like autonomous driving and robotics. Recent volumetric reconstruction methods achieve impressive results, but generally require expensive 3D ground truth or multi-view supervision. We propose to leverage pre-trained 2D diffusion models and depth prediction models to generate synthetic scene geometry from a single image. This can then be used to distill a feed-forward scene reconstruction model. Our experiments on the challenging KITTI-360 and Waymo datasets demonstrate that our method matches or outperforms state-of-the-art baselines that use multi-view supervision, and offers unique advantages, for example regarding dynamic scenes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_02323 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images Wulff, Philipp Wimbauer, Felix Muhle, Dominik Cremers, Daniel Computer Vision and Pattern Recognition Volumetric scene reconstruction from a single image is crucial for a broad range of applications like autonomous driving and robotics. Recent volumetric reconstruction methods achieve impressive results, but generally require expensive 3D ground truth or multi-view supervision. We propose to leverage pre-trained 2D diffusion models and depth prediction models to generate synthetic scene geometry from a single image. This can then be used to distill a feed-forward scene reconstruction model. Our experiments on the challenging KITTI-360 and Waymo datasets demonstrate that our method matches or outperforms state-of-the-art baselines that use multi-view supervision, and offers unique advantages, for example regarding dynamic scenes. |
| title | Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.02323 |