SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Beining, Zhu, Siting, Wang, Hesheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913944314052608
author Xu, Beining
Zhu, Siting
Wang, Hesheng
author_facet Xu, Beining
Zhu, Siting
Wang, Hesheng
contents We propose SGLoc, a novel localization system that directly regresses camera poses from 3D Gaussian Splatting (3DGS) representation by leveraging semantic information. Our method utilizes the semantic relationship between 2D image and 3D scene representation to estimate the 6DoF pose without prior pose information. In this system, we introduce a multi-level pose regression strategy that progressively estimates and refines the pose of query image from the global 3DGS map, without requiring initial pose priors. Moreover, we introduce a semantic-based global retrieval algorithm that establishes correspondences between 2D (image) and 3D (3DGS map). By matching the extracted scene semantic descriptors of 2D query image and 3DGS semantic representation, we align the image with the local region of the global 3DGS map, thereby obtaining a coarse pose estimation. Subsequently, we refine the coarse pose by iteratively optimizing the difference between the query image and the rendered image from 3DGS. Our SGLoc demonstrates superior performance over baselines on 12scenes and 7scenes datasets, showing excellent capabilities in global localization without initial pose prior. Code will be available at https://github.com/IRMVLab/SGLoc.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation
Xu, Beining
Zhu, Siting
Wang, Hesheng
Computer Vision and Pattern Recognition
Robotics
I.4.8; I.2.9
We propose SGLoc, a novel localization system that directly regresses camera poses from 3D Gaussian Splatting (3DGS) representation by leveraging semantic information. Our method utilizes the semantic relationship between 2D image and 3D scene representation to estimate the 6DoF pose without prior pose information. In this system, we introduce a multi-level pose regression strategy that progressively estimates and refines the pose of query image from the global 3DGS map, without requiring initial pose priors. Moreover, we introduce a semantic-based global retrieval algorithm that establishes correspondences between 2D (image) and 3D (3DGS map). By matching the extracted scene semantic descriptors of 2D query image and 3DGS semantic representation, we align the image with the local region of the global 3DGS map, thereby obtaining a coarse pose estimation. Subsequently, we refine the coarse pose by iteratively optimizing the difference between the query image and the rendered image from 3DGS. Our SGLoc demonstrates superior performance over baselines on 12scenes and 7scenes datasets, showing excellent capabilities in global localization without initial pose prior. Code will be available at https://github.com/IRMVLab/SGLoc.
title SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation
topic Computer Vision and Pattern Recognition
Robotics
I.4.8; I.2.9
url https://arxiv.org/abs/2507.12027