CLIP-Loc: Multi-modal Landmark Association for Global Localization in Object-based Maps

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Matsuzaki, Shigemichi, Sugino, Takuma, Tanaka, Kazuhito, Sha, Zijun, Nakaoka, Shintaro, Yoshizawa, Shintaro, Shintani, Kazuhiro
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909099323555840
author Matsuzaki, Shigemichi
Sugino, Takuma
Tanaka, Kazuhito
Sha, Zijun
Nakaoka, Shintaro
Yoshizawa, Shintaro
Shintani, Kazuhiro
author_facet Matsuzaki, Shigemichi
Sugino, Takuma
Tanaka, Kazuhito
Sha, Zijun
Nakaoka, Shintaro
Yoshizawa, Shintaro
Shintani, Kazuhiro
contents This paper describes a multi-modal data association method for global localization using object-based maps and camera images. In global localization, or relocalization, using object-based maps, existing methods typically resort to matching all possible combinations of detected objects and landmarks with the same object category, followed by inlier extraction using RANSAC or brute-force search. This approach becomes infeasible as the number of landmarks increases due to the exponential growth of correspondence candidates. In this paper, we propose labeling landmarks with natural language descriptions and extracting correspondences based on conceptual similarity with image observations using a Vision Language Model (VLM). By leveraging detailed text information, our approach efficiently extracts correspondences compared to methods using only object categories. Through experiments, we demonstrate that the proposed method enables more accurate global localization with fewer iterations compared to baseline methods, exhibiting its efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2402_06092
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLIP-Loc: Multi-modal Landmark Association for Global Localization in Object-based Maps
Matsuzaki, Shigemichi
Sugino, Takuma
Tanaka, Kazuhito
Sha, Zijun
Nakaoka, Shintaro
Yoshizawa, Shintaro
Shintani, Kazuhiro
Computer Vision and Pattern Recognition
Robotics
This paper describes a multi-modal data association method for global localization using object-based maps and camera images. In global localization, or relocalization, using object-based maps, existing methods typically resort to matching all possible combinations of detected objects and landmarks with the same object category, followed by inlier extraction using RANSAC or brute-force search. This approach becomes infeasible as the number of landmarks increases due to the exponential growth of correspondence candidates. In this paper, we propose labeling landmarks with natural language descriptions and extracting correspondences based on conceptual similarity with image observations using a Vision Language Model (VLM). By leveraging detailed text information, our approach efficiently extracts correspondences compared to methods using only object categories. Through experiments, we demonstrate that the proposed method enables more accurate global localization with fewer iterations compared to baseline methods, exhibiting its efficiency.
title CLIP-Loc: Multi-modal Landmark Association for Global Localization in Object-based Maps
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2402.06092