VISTA: Monocular Segmentation-Based Mapping for Appearance and View-Invariant Global Localization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shafferman, Hannah, Thomas, Annika, Kinnari, Jouko, Ricard, Michael, Nino, Jose, How, Jonathan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915391889997824
author Shafferman, Hannah
Thomas, Annika
Kinnari, Jouko
Ricard, Michael
Nino, Jose
How, Jonathan
author_facet Shafferman, Hannah
Thomas, Annika
Kinnari, Jouko
Ricard, Michael
Nino, Jose
How, Jonathan
contents Global localization is critical for autonomous navigation, particularly in scenarios where an agent must localize within a map generated in a different session or by another agent, as agents often have no prior knowledge about the correlation between reference frames. However, this task remains challenging in unstructured environments due to appearance changes induced by viewpoint variation, seasonal changes, spatial aliasing, and occlusions -- known failure modes for traditional place recognition methods. To address these challenges, we propose VISTA (View-Invariant Segmentation-Based Tracking for Frame Alignment), a novel open-set, monocular global localization framework that combines: 1) a front-end, object-based, segmentation and tracking pipeline, followed by 2) a submap correspondence search, which exploits geometric consistencies between environment maps to align vehicle reference frames. VISTA enables consistent localization across diverse camera viewpoints and seasonal changes, without requiring any domain-specific training or finetuning. We evaluate VISTA on seasonal and oblique-angle aerial datasets, achieving up to a 69% improvement in recall over baseline methods. Furthermore, we maintain a compact object-based map that is only 0.6% the size of the most memory-conservative baseline, making our approach capable of real-time implementation on resource-constrained platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11653
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VISTA: Monocular Segmentation-Based Mapping for Appearance and View-Invariant Global Localization
Shafferman, Hannah
Thomas, Annika
Kinnari, Jouko
Ricard, Michael
Nino, Jose
How, Jonathan
Computer Vision and Pattern Recognition
Robotics
Global localization is critical for autonomous navigation, particularly in scenarios where an agent must localize within a map generated in a different session or by another agent, as agents often have no prior knowledge about the correlation between reference frames. However, this task remains challenging in unstructured environments due to appearance changes induced by viewpoint variation, seasonal changes, spatial aliasing, and occlusions -- known failure modes for traditional place recognition methods. To address these challenges, we propose VISTA (View-Invariant Segmentation-Based Tracking for Frame Alignment), a novel open-set, monocular global localization framework that combines: 1) a front-end, object-based, segmentation and tracking pipeline, followed by 2) a submap correspondence search, which exploits geometric consistencies between environment maps to align vehicle reference frames. VISTA enables consistent localization across diverse camera viewpoints and seasonal changes, without requiring any domain-specific training or finetuning. We evaluate VISTA on seasonal and oblique-angle aerial datasets, achieving up to a 69% improvement in recall over baseline methods. Furthermore, we maintain a compact object-based map that is only 0.6% the size of the most memory-conservative baseline, making our approach capable of real-time implementation on resource-constrained platforms.
title VISTA: Monocular Segmentation-Based Mapping for Appearance and View-Invariant Global Localization
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2507.11653