LatentBKI: Open-Dictionary Continuous Mapping in Visual-Language Latent Spaces with Quantifiable Uncertainty

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wilson, Joey, Xu, Ruihan, Sun, Yile, Ewen, Parker, Zhu, Minghan, Barton, Kira, Ghaffari, Maani
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913659731574784
author Wilson, Joey
Xu, Ruihan
Sun, Yile
Ewen, Parker
Zhu, Minghan
Barton, Kira
Ghaffari, Maani
author_facet Wilson, Joey
Xu, Ruihan
Sun, Yile
Ewen, Parker
Zhu, Minghan
Barton, Kira
Ghaffari, Maani
contents This paper introduces a novel probabilistic mapping algorithm, LatentBKI, which enables open-vocabulary mapping with quantifiable uncertainty. Traditionally, semantic mapping algorithms focus on a fixed set of semantic categories which limits their applicability for complex robotic tasks. Vision-Language (VL) models have recently emerged as a technique to jointly model language and visual features in a latent space, enabling semantic recognition beyond a predefined, fixed set of semantic classes. LatentBKI recurrently incorporates neural embeddings from VL models into a voxel map with quantifiable uncertainty, leveraging the spatial correlations of nearby observations through Bayesian Kernel Inference (BKI). LatentBKI is evaluated against similar explicit semantic mapping and VL mapping frameworks on the popular Matterport3D and Semantic KITTI datasets, demonstrating that LatentBKI maintains the probabilistic benefits of continuous mapping with the additional benefit of open-dictionary queries. Real-world experiments demonstrate applicability to challenging indoor environments.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11783
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LatentBKI: Open-Dictionary Continuous Mapping in Visual-Language Latent Spaces with Quantifiable Uncertainty
Wilson, Joey
Xu, Ruihan
Sun, Yile
Ewen, Parker
Zhu, Minghan
Barton, Kira
Ghaffari, Maani
Computer Vision and Pattern Recognition
Robotics
This paper introduces a novel probabilistic mapping algorithm, LatentBKI, which enables open-vocabulary mapping with quantifiable uncertainty. Traditionally, semantic mapping algorithms focus on a fixed set of semantic categories which limits their applicability for complex robotic tasks. Vision-Language (VL) models have recently emerged as a technique to jointly model language and visual features in a latent space, enabling semantic recognition beyond a predefined, fixed set of semantic classes. LatentBKI recurrently incorporates neural embeddings from VL models into a voxel map with quantifiable uncertainty, leveraging the spatial correlations of nearby observations through Bayesian Kernel Inference (BKI). LatentBKI is evaluated against similar explicit semantic mapping and VL mapping frameworks on the popular Matterport3D and Semantic KITTI datasets, demonstrating that LatentBKI maintains the probabilistic benefits of continuous mapping with the additional benefit of open-dictionary queries. Real-world experiments demonstrate applicability to challenging indoor environments.
title LatentBKI: Open-Dictionary Continuous Mapping in Visual-Language Latent Spaces with Quantifiable Uncertainty
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2410.11783