Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zheyang, Aryal, Jagannath, Nahavandi, Saeid, Lu, Xuequan, Lim, Chee Peng, Wei, Lei, Zhou, Hailing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918032106848256
author Huang, Zheyang
Aryal, Jagannath
Nahavandi, Saeid
Lu, Xuequan
Lim, Chee Peng
Wei, Lei
Zhou, Hailing
author_facet Huang, Zheyang
Aryal, Jagannath
Nahavandi, Saeid
Lu, Xuequan
Lim, Chee Peng
Wei, Lei
Zhou, Hailing
contents Cross-view geo-localization determines the location of a query image, captured by a drone or ground-based camera, by matching it to a geo-referenced satellite image. While traditional approaches focus on image-level localization, many applications, such as search-and-rescue, infrastructure inspection, and precision delivery, demand object-level accuracy. This enables users to prompt a specific object with a single click on a drone image to retrieve precise geo-tagged information of the object. However, variations in viewpoints, timing, and imaging conditions pose significant challenges, especially when identifying visually similar objects in extensive satellite imagery. To address these challenges, we propose an Object-level Cross-view Geo-localization Network (OCGNet). It integrates user-specified click locations using Gaussian Kernel Transfer (GKT) to preserve location information throughout the network. This cue is dually embedded into the feature encoder and feature matching blocks, ensuring robust object-specific localization. Additionally, OCGNet incorporates a Location Enhancement (LE) module and a Multi-Head Cross Attention (MHCA) module to adaptively emphasize object-specific features or expand focus to relevant contextual regions when necessary. OCGNet achieves state-of-the-art performance on a public dataset, CVOGL. It also demonstrates few-shot learning capabilities, effectively generalizing from limited examples, making it suitable for diverse applications (https://github.com/ZheyangH/OCGNet).
format Preprint
id arxiv_https___arxiv_org_abs_2505_17911
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention
Huang, Zheyang
Aryal, Jagannath
Nahavandi, Saeid
Lu, Xuequan
Lim, Chee Peng
Wei, Lei
Zhou, Hailing
Computer Vision and Pattern Recognition
Artificial Intelligence
Cross-view geo-localization determines the location of a query image, captured by a drone or ground-based camera, by matching it to a geo-referenced satellite image. While traditional approaches focus on image-level localization, many applications, such as search-and-rescue, infrastructure inspection, and precision delivery, demand object-level accuracy. This enables users to prompt a specific object with a single click on a drone image to retrieve precise geo-tagged information of the object. However, variations in viewpoints, timing, and imaging conditions pose significant challenges, especially when identifying visually similar objects in extensive satellite imagery. To address these challenges, we propose an Object-level Cross-view Geo-localization Network (OCGNet). It integrates user-specified click locations using Gaussian Kernel Transfer (GKT) to preserve location information throughout the network. This cue is dually embedded into the feature encoder and feature matching blocks, ensuring robust object-specific localization. Additionally, OCGNet incorporates a Location Enhancement (LE) module and a Multi-Head Cross Attention (MHCA) module to adaptively emphasize object-specific features or expand focus to relevant contextual regions when necessary. OCGNet achieves state-of-the-art performance on a public dataset, CVOGL. It also demonstrates few-shot learning capabilities, effectively generalizing from limited examples, making it suitable for diverse applications (https://github.com/ZheyangH/OCGNet).
title Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.17911