DC-VLAQ: Query-Residual Aggregation for Robust Visual Place Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Hanyu, Zhan, Zhihao, Ming, Yuhang, Li, Liang, Hou, Dibo, Civera, Javier, Kong, Wanzeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914263327571968
author Zhu, Hanyu
Zhan, Zhihao
Ming, Yuhang
Li, Liang
Hou, Dibo
Civera, Javier
Kong, Wanzeng
author_facet Zhu, Hanyu
Zhan, Zhihao
Ming, Yuhang
Li, Liang
Hou, Dibo
Civera, Javier
Kong, Wanzeng
contents One of the central challenges in visual place recognition (VPR) is learning a robust global representation that remains discriminative under large viewpoint changes, illumination variations, and severe domain shifts. While visual foundation models (VFMs) provide strong local features, most existing methods rely on a single model, overlooking the complementary cues offered by different VFMs. However, exploiting such complementary information inevitably alters token distributions, which challenges the stability of existing query-based global aggregation schemes. To address these challenges, we propose DC-VLAQ, a representation-centric framework that integrates the fusion of complementary VFMs and robust global aggregation. Specifically, we first introduce a lightweight residual-guided complementary fusion that anchors representations in the DINOv2 feature space while injecting complementary semantics from CLIP through a learned residual correction. In addition, we propose the Vector of Local Aggregated Queries (VLAQ), a query--residual global aggregation scheme that encodes local tokens by their residual responses to learnable queries, resulting in improved stability and the preservation of fine-grained discriminative cues. Extensive experiments on standard VPR benchmarks, including Pitts30k, Tokyo24/7, MSLS, Nordland, SPED, and AmsterTime, demonstrate that DC-VLAQ consistently outperforms strong baselines and achieves state-of-the-art performance, particularly under challenging domain shifts and long-term appearance changes.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12729
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DC-VLAQ: Query-Residual Aggregation for Robust Visual Place Recognition
Zhu, Hanyu
Zhan, Zhihao
Ming, Yuhang
Li, Liang
Hou, Dibo
Civera, Javier
Kong, Wanzeng
Computer Vision and Pattern Recognition
Robotics
One of the central challenges in visual place recognition (VPR) is learning a robust global representation that remains discriminative under large viewpoint changes, illumination variations, and severe domain shifts. While visual foundation models (VFMs) provide strong local features, most existing methods rely on a single model, overlooking the complementary cues offered by different VFMs. However, exploiting such complementary information inevitably alters token distributions, which challenges the stability of existing query-based global aggregation schemes. To address these challenges, we propose DC-VLAQ, a representation-centric framework that integrates the fusion of complementary VFMs and robust global aggregation. Specifically, we first introduce a lightweight residual-guided complementary fusion that anchors representations in the DINOv2 feature space while injecting complementary semantics from CLIP through a learned residual correction. In addition, we propose the Vector of Local Aggregated Queries (VLAQ), a query--residual global aggregation scheme that encodes local tokens by their residual responses to learnable queries, resulting in improved stability and the preservation of fine-grained discriminative cues. Extensive experiments on standard VPR benchmarks, including Pitts30k, Tokyo24/7, MSLS, Nordland, SPED, and AmsterTime, demonstrate that DC-VLAQ consistently outperforms strong baselines and achieves state-of-the-art performance, particularly under challenging domain shifts and long-term appearance changes.
title DC-VLAQ: Query-Residual Aggregation for Robust Visual Place Recognition
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2601.12729