Where am I? Cross-View Geo-localization with Natural Language Descriptions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Junyan, Lin, Honglin, Ou, Leyan, Chen, Dairong, Wang, Zihao, Zhu, Qi, He, Conghui, Li, Weijia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909560143347712
author Ye, Junyan
Lin, Honglin
Ou, Leyan
Chen, Dairong
Wang, Zihao
Zhu, Qi
He, Conghui
Li, Weijia
author_facet Ye, Junyan
Lin, Honglin
Ou, Leyan
Chen, Dairong
Wang, Zihao
Zhu, Qi
He, Conghui
Li, Weijia
contents Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM. However, most existing studies focus on image-to-image retrieval, with fewer addressing text-guided retrieval, a task vital for applications like pedestrian navigation and emergency response. In this work, we introduce a novel task for cross-view geo-localization with natural language descriptions, which aims to retrieve corresponding satellite images or OSM database based on scene text descriptions. To support this task, we construct the CVG-Text dataset by collecting cross-view data from multiple cities and employing a scene text generation approach that leverages the annotation capabilities of Large Multimodal Models to produce high-quality scene text descriptions with localization details. Additionally, we propose a novel text-based retrieval localization method, CrossText2Loc, which improves recall by 10% and demonstrates excellent long-text retrieval capabilities. In terms of explainability, it not only provides similarity scores but also offers retrieval reasons. More information can be found at https://yejy53.github.io/CVG-Text/ .
format Preprint
id arxiv_https___arxiv_org_abs_2412_17007
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Where am I? Cross-View Geo-localization with Natural Language Descriptions
Ye, Junyan
Lin, Honglin
Ou, Leyan
Chen, Dairong
Wang, Zihao
Zhu, Qi
He, Conghui
Li, Weijia
Computer Vision and Pattern Recognition
Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM. However, most existing studies focus on image-to-image retrieval, with fewer addressing text-guided retrieval, a task vital for applications like pedestrian navigation and emergency response. In this work, we introduce a novel task for cross-view geo-localization with natural language descriptions, which aims to retrieve corresponding satellite images or OSM database based on scene text descriptions. To support this task, we construct the CVG-Text dataset by collecting cross-view data from multiple cities and employing a scene text generation approach that leverages the annotation capabilities of Large Multimodal Models to produce high-quality scene text descriptions with localization details. Additionally, we propose a novel text-based retrieval localization method, CrossText2Loc, which improves recall by 10% and demonstrates excellent long-text retrieval capabilities. In terms of explainability, it not only provides similarity scores but also offers retrieval reasons. More information can be found at https://yejy53.github.io/CVG-Text/ .
title Where am I? Cross-View Geo-localization with Natural Language Descriptions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.17007