OpenNav: Efficient Open Vocabulary 3D Object Detection for Smart Wheelchair Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rahman, Muhammad Rameez ur, Simonetto, Piero, Polato, Anna, Pasti, Francesco, Tonin, Luca, Vascon, Sebastiano
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910576209297408
author Rahman, Muhammad Rameez ur
Simonetto, Piero
Polato, Anna
Pasti, Francesco
Tonin, Luca
Vascon, Sebastiano
author_facet Rahman, Muhammad Rameez ur
Simonetto, Piero
Polato, Anna
Pasti, Francesco
Tonin, Luca
Vascon, Sebastiano
contents Open vocabulary 3D object detection (OV3D) allows precise and extensible object recognition crucial for adapting to diverse environments encountered in assistive robotics. This paper presents OpenNav, a zero-shot 3D object detection pipeline based on RGB-D images for smart wheelchairs. Our pipeline integrates an open-vocabulary 2D object detector with a mask generator for semantic segmentation, followed by depth isolation and point cloud construction to create 3D bounding boxes. The smart wheelchair exploits these 3D bounding boxes to identify potential targets and navigate safely. We demonstrate OpenNav's performance through experiments on the Replica dataset and we report preliminary results with a real wheelchair. OpenNav improves state-of-the-art significantly on the Replica dataset at mAP25 (+9pts) and mAP50 (+5pts) with marginal improvement at mAP. The code is publicly available at this link: https://github.com/EasyWalk-PRIN/OpenNav.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13936
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OpenNav: Efficient Open Vocabulary 3D Object Detection for Smart Wheelchair Navigation
Rahman, Muhammad Rameez ur
Simonetto, Piero
Polato, Anna
Pasti, Francesco
Tonin, Luca
Vascon, Sebastiano
Computer Vision and Pattern Recognition
Open vocabulary 3D object detection (OV3D) allows precise and extensible object recognition crucial for adapting to diverse environments encountered in assistive robotics. This paper presents OpenNav, a zero-shot 3D object detection pipeline based on RGB-D images for smart wheelchairs. Our pipeline integrates an open-vocabulary 2D object detector with a mask generator for semantic segmentation, followed by depth isolation and point cloud construction to create 3D bounding boxes. The smart wheelchair exploits these 3D bounding boxes to identify potential targets and navigate safely. We demonstrate OpenNav's performance through experiments on the Replica dataset and we report preliminary results with a real wheelchair. OpenNav improves state-of-the-art significantly on the Replica dataset at mAP25 (+9pts) and mAP50 (+5pts) with marginal improvement at mAP. The code is publicly available at this link: https://github.com/EasyWalk-PRIN/OpenNav.
title OpenNav: Efficient Open Vocabulary 3D Object Detection for Smart Wheelchair Navigation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.13936