VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jingyi, Qi, Zhangshuo, Yan, Zhongmiao, Gao, Xuyu, Jiao, Qianyun, Xia, Songpengcheng, Chen, Xieyuanli, Pei, Ling
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914344084701184
author Xu, Jingyi
Qi, Zhangshuo
Yan, Zhongmiao
Gao, Xuyu
Jiao, Qianyun
Xia, Songpengcheng
Chen, Xieyuanli
Pei, Ling
author_facet Xu, Jingyi
Qi, Zhangshuo
Yan, Zhongmiao
Gao, Xuyu
Jiao, Qianyun
Xia, Songpengcheng
Chen, Xieyuanli
Pei, Ling
contents In autonomous driving, robust place recognition is critical for global localization and loop closure detection. While inter-modality fusion of camera and LiDAR data in multimodal place recognition (MPR) has shown promise in overcoming the limitations of unimodal counterparts, existing MPR methods basically attend to hand-crafted fusion strategies and heavily parameterized backbones that require costly retraining. To address this, we propose VGGT-MPR, a multimodal place recognition framework that adopts the Visual Geometry Grounded Transformer (VGGT) as a unified geometric engine for both global retrieval and re-ranking. In the global retrieval stage, VGGT extracts geometrically-rich visual embeddings through prior depth-aware and point map supervision, and densifies sparse LiDAR point clouds with predicted depth maps to improve structural representation. This enhances the discriminative ability of fused multimodal features and produces global descriptors for fast retrieval. Beyond global retrieval, we design a training-free re-ranking mechanism that exploits VGGT's cross-view keypoint-tracking capability. By combining mask-guided keypoint extraction with confidence-aware correspondence scoring, our proposed re-ranking mechanism effectively refines retrieval results without additional parameter optimization. Extensive experiments on large-scale autonomous driving benchmarks and our self-collected data demonstrate that VGGT-MPR achieves state-of-the-art performance, exhibiting strong robustness to severe environmental changes, viewpoint shifts, and occlusions. Our code and data will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19735
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
Xu, Jingyi
Qi, Zhangshuo
Yan, Zhongmiao
Gao, Xuyu
Jiao, Qianyun
Xia, Songpengcheng
Chen, Xieyuanli
Pei, Ling
Computer Vision and Pattern Recognition
In autonomous driving, robust place recognition is critical for global localization and loop closure detection. While inter-modality fusion of camera and LiDAR data in multimodal place recognition (MPR) has shown promise in overcoming the limitations of unimodal counterparts, existing MPR methods basically attend to hand-crafted fusion strategies and heavily parameterized backbones that require costly retraining. To address this, we propose VGGT-MPR, a multimodal place recognition framework that adopts the Visual Geometry Grounded Transformer (VGGT) as a unified geometric engine for both global retrieval and re-ranking. In the global retrieval stage, VGGT extracts geometrically-rich visual embeddings through prior depth-aware and point map supervision, and densifies sparse LiDAR point clouds with predicted depth maps to improve structural representation. This enhances the discriminative ability of fused multimodal features and produces global descriptors for fast retrieval. Beyond global retrieval, we design a training-free re-ranking mechanism that exploits VGGT's cross-view keypoint-tracking capability. By combining mask-guided keypoint extraction with confidence-aware correspondence scoring, our proposed re-ranking mechanism effectively refines retrieval results without additional parameter optimization. Extensive experiments on large-scale autonomous driving benchmarks and our self-collected data demonstrate that VGGT-MPR achieves state-of-the-art performance, exhibiting strong robustness to severe environmental changes, viewpoint shifts, and occlusions. Our code and data will be made publicly available.
title VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.19735