Where am I? Cross-View Geo-localization with Natural Language Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909560143347712 |
|---|---|
| author | Ye, Junyan Lin, Honglin Ou, Leyan Chen, Dairong Wang, Zihao Zhu, Qi He, Conghui Li, Weijia |
| author_facet | Ye, Junyan Lin, Honglin Ou, Leyan Chen, Dairong Wang, Zihao Zhu, Qi He, Conghui Li, Weijia |
| contents | Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM. However, most existing studies focus on image-to-image retrieval, with fewer addressing text-guided retrieval, a task vital for applications like pedestrian navigation and emergency response. In this work, we introduce a novel task for cross-view geo-localization with natural language descriptions, which aims to retrieve corresponding satellite images or OSM database based on scene text descriptions. To support this task, we construct the CVG-Text dataset by collecting cross-view data from multiple cities and employing a scene text generation approach that leverages the annotation capabilities of Large Multimodal Models to produce high-quality scene text descriptions with localization details. Additionally, we propose a novel text-based retrieval localization method, CrossText2Loc, which improves recall by 10% and demonstrates excellent long-text retrieval capabilities. In terms of explainability, it not only provides similarity scores but also offers retrieval reasons. More information can be found at https://yejy53.github.io/CVG-Text/ . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_17007 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Where am I? Cross-View Geo-localization with Natural Language Descriptions Ye, Junyan Lin, Honglin Ou, Leyan Chen, Dairong Wang, Zihao Zhu, Qi He, Conghui Li, Weijia Computer Vision and Pattern Recognition Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM. However, most existing studies focus on image-to-image retrieval, with fewer addressing text-guided retrieval, a task vital for applications like pedestrian navigation and emergency response. In this work, we introduce a novel task for cross-view geo-localization with natural language descriptions, which aims to retrieve corresponding satellite images or OSM database based on scene text descriptions. To support this task, we construct the CVG-Text dataset by collecting cross-view data from multiple cities and employing a scene text generation approach that leverages the annotation capabilities of Large Multimodal Models to produce high-quality scene text descriptions with localization details. Additionally, we propose a novel text-based retrieval localization method, CrossText2Loc, which improves recall by 10% and demonstrates excellent long-text retrieval capabilities. In terms of explainability, it not only provides similarity scores but also offers retrieval reasons. More information can be found at https://yejy53.github.io/CVG-Text/ . |
| title | Where am I? Cross-View Geo-localization with Natural Language Descriptions |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.17007 |