Saved in:
Bibliographic Details
Main Authors: Li, Hongyu, Padir, Taskin, Jiang, Huaizu
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.12039
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913270105899008
author Li, Hongyu
Padir, Taskin
Jiang, Huaizu
author_facet Li, Hongyu
Padir, Taskin
Jiang, Huaizu
contents Visual navigation has received significant attention recently. Most of the prior works focus on predicting navigation actions based on semantic features extracted from visual encoders. However, these approaches often rely on large datasets and exhibit limited generalizability. In contrast, our approach draws inspiration from traditional navigation planners that operate on geometric representations, such as occupancy maps. We propose StereoNavNet (SNN), a novel visual navigation approach employing a modular learning framework comprising perception and policy modules. Within the perception module, we estimate an auxiliary 3D voxel occupancy grid from stereo RGB images and extract geometric features from it. These features, along with user-defined goals, are utilized by the policy module to predict navigation actions. Through extensive empirical evaluation, we demonstrate that SNN outperforms baseline approaches in terms of success rates, success weighted by path length, and navigation error. Furthermore, SNN exhibits better generalizability, characterized by maintaining leading performance when navigating across previously unseen environments.
format Preprint
id arxiv_https___arxiv_org_abs_2403_12039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StereoNavNet: Learning to Navigate using Stereo Cameras with Auxiliary Occupancy Voxels
Li, Hongyu
Padir, Taskin
Jiang, Huaizu
Robotics
Visual navigation has received significant attention recently. Most of the prior works focus on predicting navigation actions based on semantic features extracted from visual encoders. However, these approaches often rely on large datasets and exhibit limited generalizability. In contrast, our approach draws inspiration from traditional navigation planners that operate on geometric representations, such as occupancy maps. We propose StereoNavNet (SNN), a novel visual navigation approach employing a modular learning framework comprising perception and policy modules. Within the perception module, we estimate an auxiliary 3D voxel occupancy grid from stereo RGB images and extract geometric features from it. These features, along with user-defined goals, are utilized by the policy module to predict navigation actions. Through extensive empirical evaluation, we demonstrate that SNN outperforms baseline approaches in terms of success rates, success weighted by path length, and navigation error. Furthermore, SNN exhibits better generalizability, characterized by maintaining leading performance when navigating across previously unseen environments.
title StereoNavNet: Learning to Navigate using Stereo Cameras with Auxiliary Occupancy Voxels
topic Robotics
url https://arxiv.org/abs/2403.12039