GRLoc: Geometric Representation Regression for Visual Localization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Changyang, Ma, Xuejian, Liu, Lixiang, Li, Zhan, Yan, Qingan, Xu, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914578765447168
author Li, Changyang
Ma, Xuejian
Liu, Lixiang
Li, Zhan
Yan, Qingan
Xu, Yi
author_facet Li, Changyang
Ma, Xuejian
Liu, Lixiang
Li, Zhan
Yan, Qingan
Xu, Yi
contents Absolute Pose Regression (APR) has emerged as a compelling paradigm for visual localization. However, APR models typically operate as black boxes, directly regressing a 6-DoF pose from a query image, which can lead to memorizing training views rather than understanding 3D scene geometry. In this work, we propose a geometrically-grounded alternative. Inspired by novel view synthesis, which renders images from intermediate geometric representations, we reformulate APR as its inverse that regresses the underlying 3D representations directly from the image, and we name this paradigm Geometric Representation Regression (GRR). Our model explicitly predicts two disentangled geometric representations in the world coordinate system: (1) a raymap's directions to estimate camera rotation, and (2) a corresponding pointmap to estimate camera translation. The final camera pose is then recovered from these geometric components using a differentiable deterministic solver. This disentangled approach, which separates the learned visual-to-geometry mapping from the final pose calculation, introduces a strong geometric prior into the network. We find that the explicit decoupling of rotation and translation predictions measurably boosts performance. We demonstrate state-of-the-art performance on 7-Scenes and Cambridge Landmarks datasets, validating that modeling the inverse rendering process is a more robust path toward generalizable absolute pose estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRLoc: Geometric Representation Regression for Visual Localization
Li, Changyang
Ma, Xuejian
Liu, Lixiang
Li, Zhan
Yan, Qingan
Xu, Yi
Computer Vision and Pattern Recognition
Absolute Pose Regression (APR) has emerged as a compelling paradigm for visual localization. However, APR models typically operate as black boxes, directly regressing a 6-DoF pose from a query image, which can lead to memorizing training views rather than understanding 3D scene geometry. In this work, we propose a geometrically-grounded alternative. Inspired by novel view synthesis, which renders images from intermediate geometric representations, we reformulate APR as its inverse that regresses the underlying 3D representations directly from the image, and we name this paradigm Geometric Representation Regression (GRR). Our model explicitly predicts two disentangled geometric representations in the world coordinate system: (1) a raymap's directions to estimate camera rotation, and (2) a corresponding pointmap to estimate camera translation. The final camera pose is then recovered from these geometric components using a differentiable deterministic solver. This disentangled approach, which separates the learned visual-to-geometry mapping from the final pose calculation, introduces a strong geometric prior into the network. We find that the explicit decoupling of rotation and translation predictions measurably boosts performance. We demonstrate state-of-the-art performance on 7-Scenes and Cambridge Landmarks datasets, validating that modeling the inverse rendering process is a more robust path toward generalizable absolute pose estimation.
title GRLoc: Geometric Representation Regression for Visual Localization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.13864