Pre-Trained Masked Image Model for Mobile Robot Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sharma, Vishnu Dutt, Singh, Anukriti, Tokekar, Pratap
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910382431404032
author Sharma, Vishnu Dutt
Singh, Anukriti
Tokekar, Pratap
author_facet Sharma, Vishnu Dutt
Singh, Anukriti
Tokekar, Pratap
contents 2D top-down maps are commonly used for the navigation and exploration of mobile robots through unknown areas. Typically, the robot builds the navigation maps incrementally from local observations using onboard sensors. Recent works have shown that predicting the structural patterns in the environment through learning-based approaches can greatly enhance task efficiency. While many such works build task-specific networks using limited datasets, we show that the existing foundational vision networks can accomplish the same without any fine-tuning. Specifically, we use Masked Autoencoders, pre-trained on street images, to present novel applications for field-of-view expansion, single-agent topological exploration, and multi-agent exploration for indoor mapping, across different input modalities. Our work motivates the use of foundational vision models for generalized structure prediction-driven applications, especially in the dearth of training data. For more qualitative results see https://raaslab.org/projects/MIM4Robots.
format Preprint
id arxiv_https___arxiv_org_abs_2310_07021
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Pre-Trained Masked Image Model for Mobile Robot Navigation
Sharma, Vishnu Dutt
Singh, Anukriti
Tokekar, Pratap
Robotics
Computer Vision and Pattern Recognition
2D top-down maps are commonly used for the navigation and exploration of mobile robots through unknown areas. Typically, the robot builds the navigation maps incrementally from local observations using onboard sensors. Recent works have shown that predicting the structural patterns in the environment through learning-based approaches can greatly enhance task efficiency. While many such works build task-specific networks using limited datasets, we show that the existing foundational vision networks can accomplish the same without any fine-tuning. Specifically, we use Masked Autoencoders, pre-trained on street images, to present novel applications for field-of-view expansion, single-agent topological exploration, and multi-agent exploration for indoor mapping, across different input modalities. Our work motivates the use of foundational vision models for generalized structure prediction-driven applications, especially in the dearth of training data. For more qualitative results see https://raaslab.org/projects/MIM4Robots.
title Pre-Trained Masked Image Model for Mobile Robot Navigation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.07021