FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Laina, Sebastián Barbas, Boche, Simon, Papatheodorou, Sotiris, Schaefer, Simon, Jung, Jaehyung, Oleynikova, Helen, Leutenegger, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908869321555968
author Laina, Sebastián Barbas
Boche, Simon
Papatheodorou, Sotiris
Schaefer, Simon
Jung, Jaehyung
Oleynikova, Helen
Leutenegger, Stefan
author_facet Laina, Sebastián Barbas
Boche, Simon
Papatheodorou, Sotiris
Schaefer, Simon
Jung, Jaehyung
Oleynikova, Helen
Leutenegger, Stefan
contents Geometrically accurate and semantically expressive map representations have proven invaluable for robot deployment and task planning in unknown environments. Nevertheless, real-time, open-vocabulary semantic understanding of large-scale unknown environments still presents open challenges, mainly due to computational requirements. In this paper we present FindAnything, an open-world mapping framework that incorporates vision-language information into dense volumetric submaps. Thanks to the use of vision-language features, FindAnything combines pure geometric and open-vocabulary semantic information for a higher level of understanding. It proposes an efficient storage of open-vocabulary information through the aggregation of features at the object level. Pixelwise vision-language features are aggregated based on eSAM segments, which are in turn integrated into object-centric volumetric submaps, providing a mapping from open-vocabulary queries to 3D geometry that is scalable also in terms of memory usage. We demonstrate that FindAnything performs on par with the state-of-the-art in terms of semantic accuracy while being substantially faster and more memory-efficient, allowing its deployment in large-scale environments and on resourceconstrained devices, such as MAVs. We show that the real-time capabilities of FindAnything make it useful for downstream tasks, such as autonomous MAV exploration in a simulated Search and Rescue scenario. Project Page: https://ethz-mrl.github.io/findanything/.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08603
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment
Laina, Sebastián Barbas
Boche, Simon
Papatheodorou, Sotiris
Schaefer, Simon
Jung, Jaehyung
Oleynikova, Helen
Leutenegger, Stefan
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Geometrically accurate and semantically expressive map representations have proven invaluable for robot deployment and task planning in unknown environments. Nevertheless, real-time, open-vocabulary semantic understanding of large-scale unknown environments still presents open challenges, mainly due to computational requirements. In this paper we present FindAnything, an open-world mapping framework that incorporates vision-language information into dense volumetric submaps. Thanks to the use of vision-language features, FindAnything combines pure geometric and open-vocabulary semantic information for a higher level of understanding. It proposes an efficient storage of open-vocabulary information through the aggregation of features at the object level. Pixelwise vision-language features are aggregated based on eSAM segments, which are in turn integrated into object-centric volumetric submaps, providing a mapping from open-vocabulary queries to 3D geometry that is scalable also in terms of memory usage. We demonstrate that FindAnything performs on par with the state-of-the-art in terms of semantic accuracy while being substantially faster and more memory-efficient, allowing its deployment in large-scale environments and on resourceconstrained devices, such as MAVs. We show that the real-time capabilities of FindAnything make it useful for downstream tasks, such as autonomous MAV exploration in a simulated Search and Rescue scenario. Project Page: https://ethz-mrl.github.io/findanything/.
title FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.08603