OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous Driving

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Heng, Duan, Yifan, Zhang, Xinran, Liu, Haiyi, Ji, Jianmin, Zhang, Yanyong
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910382408335360
author Li, Heng
Duan, Yifan
Zhang, Xinran
Liu, Haiyi
Ji, Jianmin
Zhang, Yanyong
author_facet Li, Heng
Duan, Yifan
Zhang, Xinran
Liu, Haiyi
Ji, Jianmin
Zhang, Yanyong
contents Visual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupancy, thereby circumventing the traditional need for concurrent estimation of ego poses and landmark locations. Within this framework, we utilize the TPV-Former to convert surround view cameras' images into 3D semantic occupancy. Addressing the challenges presented by this transformation, we have specifically tailored a pose estimation and mapping algorithm that incorporates Semantic Label Filter, Dynamic Object Filter, and finally, utilizes Voxel PFilter for maintaining a consistent global semantic map. Evaluations on the Occ3D-nuScenes not only showcase a 20.6% improvement in Success Ratio and a 29.6% enhancement in trajectory accuracy against ORB-SLAM3, but also emphasize our ability to construct a comprehensive map. Our implementation is open-sourced and available at: https://github.com/USTCLH/OCC-VO.
format Preprint
id arxiv_https___arxiv_org_abs_2309_11011
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous Driving
Li, Heng
Duan, Yifan
Zhang, Xinran
Liu, Haiyi
Ji, Jianmin
Zhang, Yanyong
Robotics
Visual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupancy, thereby circumventing the traditional need for concurrent estimation of ego poses and landmark locations. Within this framework, we utilize the TPV-Former to convert surround view cameras' images into 3D semantic occupancy. Addressing the challenges presented by this transformation, we have specifically tailored a pose estimation and mapping algorithm that incorporates Semantic Label Filter, Dynamic Object Filter, and finally, utilizes Voxel PFilter for maintaining a consistent global semantic map. Evaluations on the Occ3D-nuScenes not only showcase a 20.6% improvement in Success Ratio and a 29.6% enhancement in trajectory accuracy against ORB-SLAM3, but also emphasize our ability to construct a comprehensive map. Our implementation is open-sourced and available at: https://github.com/USTCLH/OCC-VO.
title OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous Driving
topic Robotics
url https://arxiv.org/abs/2309.11011