SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Bowen, Jia, Pengyue, Wang, Wanyu, Xu, Derong, Cheng, Jiawei, Dong, Jiancheng, Han, Xiao, Zhao, Zimo, Zhang, Chao, Yu, Bowen, Hong, Fangyu, Zhao, Xiangyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910223109718016
author Liu, Bowen
Jia, Pengyue
Wang, Wanyu
Xu, Derong
Cheng, Jiawei
Dong, Jiancheng
Han, Xiao
Zhao, Zimo
Zhang, Chao
Yu, Bowen
Hong, Fangyu
Zhao, Xiangyu
author_facet Liu, Bowen
Jia, Pengyue
Wang, Wanyu
Xu, Derong
Cheng, Jiawei
Dong, Jiancheng
Han, Xiao
Zhao, Zimo
Zhang, Chao
Yu, Bowen
Hong, Fangyu
Zhao, Xiangyu
contents Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Unmanned Aerial Vehicle (UAV) queries by matching them against an extensive geo-tagged satellite image database. Most existing methods learn separate feature representations for each view and determine the final prediction using naive heuristics to assess feature similarity, thereby neglecting to model the crucial cross-view relationships. In this paper, we propose SkyLink, a novel plug-and-play ranking framework that pioneers joint relational modeling of inter-view relationships to enhance cross-view UAV geolocalization. SkyLink leverages a Large Vision-Language Model (LVLM) to model the intricate visual-semantic relationships between UAV and satellite views, facilitating effective cross-view matching. To further refine the learning process, we introduce a relational-aware loss. It leverages soft labels to provide a more nuanced supervision signal, mitigating the harsh penalty on near-positive pairs. This approach enhances both training stability and the model's discriminative capacity. Extensive experiments conducted across multiple base retrieval architectures and benchmark datasets demonstrate that SkyLink significantly boosts the ranking effectiveness of existing models, consistently achieving superior performance in various challenging scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08063
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization
Liu, Bowen
Jia, Pengyue
Wang, Wanyu
Xu, Derong
Cheng, Jiawei
Dong, Jiancheng
Han, Xiao
Zhao, Zimo
Zhang, Chao
Yu, Bowen
Hong, Fangyu
Zhao, Xiangyu
Computer Vision and Pattern Recognition
Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Unmanned Aerial Vehicle (UAV) queries by matching them against an extensive geo-tagged satellite image database. Most existing methods learn separate feature representations for each view and determine the final prediction using naive heuristics to assess feature similarity, thereby neglecting to model the crucial cross-view relationships. In this paper, we propose SkyLink, a novel plug-and-play ranking framework that pioneers joint relational modeling of inter-view relationships to enhance cross-view UAV geolocalization. SkyLink leverages a Large Vision-Language Model (LVLM) to model the intricate visual-semantic relationships between UAV and satellite views, facilitating effective cross-view matching. To further refine the learning process, we introduce a relational-aware loss. It leverages soft labels to provide a more nuanced supervision signal, mitigating the harsh penalty on near-positive pairs. This approach enhances both training stability and the model's discriminative capacity. Extensive experiments conducted across multiple base retrieval architectures and benchmark datasets demonstrate that SkyLink significantly boosts the ranking effectiveness of existing models, consistently achieving superior performance in various challenging scenarios.
title SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.08063