Learning 3D Reconstruction with Priors in Test Time

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Lei, Wu, Haoyu, Dave, Akshat, Samaras, Dimitris
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908937271377920
author Zhou, Lei
Wu, Haoyu
Dave, Akshat
Samaras, Dimitris
author_facet Zhou, Lei
Wu, Haoyu
Dave, Akshat
Samaras, Dimitris
contents We introduce a test-time framework for multiview Transformers (MVTs) that incorporates priors (e.g., camera poses, intrinsics, and depth) to improve 3D tasks without retraining or modifying pre-trained image-only networks. Rather than feeding priors into the architecture, we cast them as constraints on the predictions and optimize the network at inference time. The optimization loss consists of a self-supervised objective and prior penalty terms. The self-supervised objective captures the compatibility among multi-view predictions and is implemented using photometric or geometric loss between renderings from other views and each view itself. Any available priors are converted into penalty terms on the corresponding output modalities. Across a series of 3D vision benchmarks, including point map estimation and camera pose estimation, our method consistently improves performance over base MVTs by a large margin. On the ETH3D, 7-Scenes, and NRGBD datasets, our method reduces the point-map distance error by more than half compared with the base image-only models. Our method also outperforms retrained prior-aware feed-forward methods, demonstrating the effectiveness of our test-time constrained optimization (TCO) framework for incorporating priors into 3D vision tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03878
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning 3D Reconstruction with Priors in Test Time
Zhou, Lei
Wu, Haoyu
Dave, Akshat
Samaras, Dimitris
Computer Vision and Pattern Recognition
We introduce a test-time framework for multiview Transformers (MVTs) that incorporates priors (e.g., camera poses, intrinsics, and depth) to improve 3D tasks without retraining or modifying pre-trained image-only networks. Rather than feeding priors into the architecture, we cast them as constraints on the predictions and optimize the network at inference time. The optimization loss consists of a self-supervised objective and prior penalty terms. The self-supervised objective captures the compatibility among multi-view predictions and is implemented using photometric or geometric loss between renderings from other views and each view itself. Any available priors are converted into penalty terms on the corresponding output modalities. Across a series of 3D vision benchmarks, including point map estimation and camera pose estimation, our method consistently improves performance over base MVTs by a large margin. On the ETH3D, 7-Scenes, and NRGBD datasets, our method reduces the point-map distance error by more than half compared with the base image-only models. Our method also outperforms retrained prior-aware feed-forward methods, demonstrating the effectiveness of our test-time constrained optimization (TCO) framework for incorporating priors into 3D vision tasks.
title Learning 3D Reconstruction with Priors in Test Time
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.03878