ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Weidong, Luo, Lun, Ye, Nanfei, Ren, Yi, Du, Shaoyi, Wang, Minhang, Xu, Jintao, Ai, Rui, Gu, Weihao, Chen, Xieyuanli
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916180442218496
author Xie, Weidong
Luo, Lun
Ye, Nanfei
Ren, Yi
Du, Shaoyi
Wang, Minhang
Xu, Jintao
Ai, Rui
Gu, Weihao
Chen, Xieyuanli
author_facet Xie, Weidong
Luo, Lun
Ye, Nanfei
Ren, Yi
Du, Shaoyi
Wang, Minhang
Xu, Jintao
Ai, Rui
Gu, Weihao
Chen, Xieyuanli
contents Place recognition is an important task for robots and autonomous cars to localize themselves and close loops in pre-built maps. While single-modal sensor-based methods have shown satisfactory performance, cross-modal place recognition that retrieving images from a point-cloud database remains a challenging problem. Current cross-modal methods transform images into 3D points using depth estimation for modality conversion, which are usually computationally intensive and need expensive labeled data for depth supervision. In this work, we introduce a fast and lightweight framework to encode images and point clouds into place-distinctive descriptors. We propose an effective Field of View (FoV) transformation module to convert point clouds into an analogous modality as images. This module eliminates the necessity for depth estimation and helps subsequent modules achieve real-time performance. We further design a non-negative factorization-based encoder to extract mutually consistent semantic features between point clouds and images. This encoder yields more distinctive global descriptors for retrieval. Experimental results on the KITTI dataset show that our proposed methods achieve state-of-the-art performance while running in real time. Additional evaluation on the HAOMO dataset covering a 17 km trajectory further shows the practical generalization capabilities. We have released the implementation of our methods as open source at: https://github.com/haomo-ai/ModaLink.git.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18762
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
Xie, Weidong
Luo, Lun
Ye, Nanfei
Ren, Yi
Du, Shaoyi
Wang, Minhang
Xu, Jintao
Ai, Rui
Gu, Weihao
Chen, Xieyuanli
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Place recognition is an important task for robots and autonomous cars to localize themselves and close loops in pre-built maps. While single-modal sensor-based methods have shown satisfactory performance, cross-modal place recognition that retrieving images from a point-cloud database remains a challenging problem. Current cross-modal methods transform images into 3D points using depth estimation for modality conversion, which are usually computationally intensive and need expensive labeled data for depth supervision. In this work, we introduce a fast and lightweight framework to encode images and point clouds into place-distinctive descriptors. We propose an effective Field of View (FoV) transformation module to convert point clouds into an analogous modality as images. This module eliminates the necessity for depth estimation and helps subsequent modules achieve real-time performance. We further design a non-negative factorization-based encoder to extract mutually consistent semantic features between point clouds and images. This encoder yields more distinctive global descriptors for retrieval. Experimental results on the KITTI dataset show that our proposed methods achieve state-of-the-art performance while running in real time. Additional evaluation on the HAOMO dataset covering a 17 km trajectory further shows the practical generalization capabilities. We have released the implementation of our methods as open source at: https://github.com/haomo-ai/ModaLink.git.
title ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2403.18762