Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Changwei, Chen, Shunpeng, Song, Yukun, Xu, Rongtao, Zhang, Zherui, Zhang, Jiguang, Yang, Haoran, Zhang, Yu, Fu, Kexue, Du, Shide, Xu, Zhiwei, Gao, Longxiang, Guo, Li, Xu, Shibiao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916688424861696
author Wang, Changwei
Chen, Shunpeng
Song, Yukun
Xu, Rongtao
Zhang, Zherui
Zhang, Jiguang
Yang, Haoran
Zhang, Yu
Fu, Kexue
Du, Shide
Xu, Zhiwei
Gao, Longxiang
Guo, Li
Xu, Shibiao
author_facet Wang, Changwei
Chen, Shunpeng
Song, Yukun
Xu, Rongtao
Zhang, Zherui
Zhang, Jiguang
Yang, Haoran
Zhang, Yu
Fu, Kexue
Du, Shide
Xu, Zhiwei
Gao, Longxiang
Guo, Li
Xu, Shibiao
contents Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency. Code and models are available at https://github.com/chenshunpeng/FoL
format Preprint
id arxiv_https___arxiv_org_abs_2504_09881
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
Wang, Changwei
Chen, Shunpeng
Song, Yukun
Xu, Rongtao
Zhang, Zherui
Zhang, Jiguang
Yang, Haoran
Zhang, Yu
Fu, Kexue
Du, Shide
Xu, Zhiwei
Gao, Longxiang
Guo, Li
Xu, Shibiao
Computer Vision and Pattern Recognition
Visual Place Recognition (VPR) is aimed at predicting the location of a query image by referencing a database of geotagged images. For VPR task, often fewer discriminative local regions in an image produce important effects while mundane background regions do not contribute or even cause perceptual aliasing because of easy overlap. However, existing methods lack precisely modeling and full exploitation of these discriminative regions. In this paper, we propose the Focus on Local (FoL) approach to stimulate the performance of image retrieval and re-ranking in VPR simultaneously by mining and exploiting reliable discriminative local regions in images and introducing pseudo-correlation supervision. First, we design two losses, Extraction-Aggregation Spatial Alignment Loss (SAL) and Foreground-Background Contrast Enhancement Loss (CEL), to explicitly model reliable discriminative local regions and use them to guide the generation of global representations and efficient re-ranking. Second, we introduce a weakly-supervised local feature training strategy based on pseudo-correspondences obtained from aggregating global features to alleviate the lack of local correspondences ground truth for the VPR task. Third, we suggest an efficient re-ranking pipeline that is efficiently and precisely based on discriminative region guidance. Finally, experimental results show that our FoL achieves the state-of-the-art on multiple VPR benchmarks in both image retrieval and re-ranking stages and also significantly outperforms existing two-stage VPR methods in terms of computational efficiency. Code and models are available at https://github.com/chenshunpeng/FoL
title Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.09881