Representing 3D sparse map points and lines for camera relocalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bui, Bach-Thuan, Bui, Huy-Hoang, Tran, Dinh-Tuan, Lee, Joo-Ho
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929258365976576
author Bui, Bach-Thuan
Bui, Huy-Hoang
Tran, Dinh-Tuan
Lee, Joo-Ho
author_facet Bui, Bach-Thuan
Bui, Huy-Hoang
Tran, Dinh-Tuan
Lee, Joo-Ho
contents Recent advancements in visual localization and mapping have demonstrated considerable success in integrating point and line features. However, expanding the localization framework to include additional mapping components frequently results in increased demand for memory and computational resources dedicated to matching tasks. In this study, we show how a lightweight neural network can learn to represent both 3D point and line features, and exhibit leading pose accuracy by harnessing the power of multiple learned mappings. Specifically, we utilize a single transformer block to encode line features, effectively transforming them into distinctive point-like descriptors. Subsequently, we treat these point and line descriptor sets as distinct yet interconnected feature sets. Through the integration of self- and cross-attention within several graph layers, our method effectively refines each feature before regressing 3D maps using two simple MLPs. In comprehensive experiments, our indoor localization findings surpass those of Hloc and Limap across both point-based and line-assisted configurations. Moreover, in outdoor scenarios, our method secures a significant lead, marking the most considerable enhancement over state-of-the-art learning-based methodologies. The source code and demo videos of this work are publicly available at: https://thpjp.github.io/pl2map/
format Preprint
id arxiv_https___arxiv_org_abs_2402_18011
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Representing 3D sparse map points and lines for camera relocalization
Bui, Bach-Thuan
Bui, Huy-Hoang
Tran, Dinh-Tuan
Lee, Joo-Ho
Computer Vision and Pattern Recognition
Recent advancements in visual localization and mapping have demonstrated considerable success in integrating point and line features. However, expanding the localization framework to include additional mapping components frequently results in increased demand for memory and computational resources dedicated to matching tasks. In this study, we show how a lightweight neural network can learn to represent both 3D point and line features, and exhibit leading pose accuracy by harnessing the power of multiple learned mappings. Specifically, we utilize a single transformer block to encode line features, effectively transforming them into distinctive point-like descriptors. Subsequently, we treat these point and line descriptor sets as distinct yet interconnected feature sets. Through the integration of self- and cross-attention within several graph layers, our method effectively refines each feature before regressing 3D maps using two simple MLPs. In comprehensive experiments, our indoor localization findings surpass those of Hloc and Limap across both point-based and line-assisted configurations. Moreover, in outdoor scenarios, our method secures a significant lead, marking the most considerable enhancement over state-of-the-art learning-based methodologies. The source code and demo videos of this work are publicly available at: https://thpjp.github.io/pl2map/
title Representing 3D sparse map points and lines for camera relocalization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.18011