Anchor-free Cross-view Object Geo-localization with Gaussian Position Encoding and Cross-view Association

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ling, Xingtao, Fu, Chenlin, Zhu, Yingying
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911185441390592
author Ling, Xingtao
Fu, Chenlin
Zhu, Yingying
author_facet Ling, Xingtao
Fu, Chenlin
Zhu, Yingying
contents Most existing cross-view object geo-localization approaches adopt anchor-based paradigm. Although effective, such methods are inherently constrained by predefined anchors. To eliminate this dependency, we first propose an anchor-free formulation for cross-view object geo-localization, termed AFGeo. AFGeo directly predicts the four directional offsets (left, right, top, bottom) to the ground-truth box for each pixel, thereby localizing the object without any predefined anchors. To obtain a more robust spatial prior, AFGeo incorporates Gaussian Position Encoding (GPE) to model the click point in the query image, mitigating the uncertainty of object position that challenges object localization in cross-view scenarios. In addition, AFGeo incorporates a Cross-view Object Association Module (CVOAM) that relates the same object and its surrounding context across viewpoints, enabling reliable localization under large cross-view appearance gaps. By adopting an anchor-free localization paradigm that integrates GPE and CVOAM with minimal parameter overhead, our model is both lightweight and computationally efficient, achieving state-of-the-art performance on benchmark datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Anchor-free Cross-view Object Geo-localization with Gaussian Position Encoding and Cross-view Association
Ling, Xingtao
Fu, Chenlin
Zhu, Yingying
Computer Vision and Pattern Recognition
Most existing cross-view object geo-localization approaches adopt anchor-based paradigm. Although effective, such methods are inherently constrained by predefined anchors. To eliminate this dependency, we first propose an anchor-free formulation for cross-view object geo-localization, termed AFGeo. AFGeo directly predicts the four directional offsets (left, right, top, bottom) to the ground-truth box for each pixel, thereby localizing the object without any predefined anchors. To obtain a more robust spatial prior, AFGeo incorporates Gaussian Position Encoding (GPE) to model the click point in the query image, mitigating the uncertainty of object position that challenges object localization in cross-view scenarios. In addition, AFGeo incorporates a Cross-view Object Association Module (CVOAM) that relates the same object and its surrounding context across viewpoints, enabling reliable localization under large cross-view appearance gaps. By adopting an anchor-free localization paradigm that integrates GPE and CVOAM with minimal parameter overhead, our model is both lightweight and computationally efficient, achieving state-of-the-art performance on benchmark datasets.
title Anchor-free Cross-view Object Geo-localization with Gaussian Position Encoding and Cross-view Association
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.25623