$R^3$: 3D Reconstruction via Relative Regression

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Congrong, Gao, Huachen, Chen, Xingyu, Xiu, Yuliang, Gao, Jun, Chen, Anpei
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910270647959552
author Xu, Congrong
Gao, Huachen
Chen, Xingyu
Xiu, Yuliang
Gao, Jun
Chen, Anpei
author_facet Xu, Congrong
Gao, Huachen
Chen, Xingyu
Xiu, Yuliang
Gao, Jun
Chen, Anpei
contents Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This dependency becomes a significant bottleneck for long-context and streaming reconstruction, as it forces the network to maintain an arbitrary temporal origin and handle translation magnitudes that grow unbounded over time. Our solution, which we call $R^3$, employs relative regression. We employ a lightweight MLP to predict confidence-weighted relative constraints. These confidences serve as a unified anchor: weighting losses during training and guiding pose aggregation during inference. $R^3$ supports both full-context offline reconstruction and causal, bounded-memory streaming. Our evaluation in both offline and streaming settings validates the effectiveness of our relative mechanism. Project page: https://kevinxu02.github.io/r3-site
format Preprint
id arxiv_https___arxiv_org_abs_2605_26519
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle $R^3$: 3D Reconstruction via Relative Regression
Xu, Congrong
Gao, Huachen
Chen, Xingyu
Xiu, Yuliang
Gao, Jun
Chen, Anpei
Computer Vision and Pattern Recognition
Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This dependency becomes a significant bottleneck for long-context and streaming reconstruction, as it forces the network to maintain an arbitrary temporal origin and handle translation magnitudes that grow unbounded over time. Our solution, which we call $R^3$, employs relative regression. We employ a lightweight MLP to predict confidence-weighted relative constraints. These confidences serve as a unified anchor: weighting losses during training and guiding pose aggregation during inference. $R^3$ supports both full-context offline reconstruction and causal, bounded-memory streaming. Our evaluation in both offline and streaming settings validates the effectiveness of our relative mechanism. Project page: https://kevinxu02.github.io/r3-site
title $R^3$: 3D Reconstruction via Relative Regression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.26519