Fully Sparse 3D Occupancy Prediction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Haisong, Chen, Yang, Wang, Haiguang, Yang, Zetong, Li, Tianyu, Zeng, Jia, Chen, Li, Li, Hongyang, Wang, Limin
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913436144762880
author Liu, Haisong
Chen, Yang
Wang, Haiguang
Yang, Zetong
Li, Tianyu
Zeng, Jia
Chen, Li
Li, Hongyang
Wang, Limin
author_facet Liu, Haisong
Chen, Yang
Wang, Haiguang
Yang, Zetong
Li, Tianyu
Zeng, Jia
Chen, Li
Li, Hongyang
Wang, Limin
contents Occupancy prediction plays a pivotal role in autonomous driving. Previous methods typically construct dense 3D volumes, neglecting the inherent sparsity of the scene and suffering from high computational costs. To bridge the gap, we introduce a novel fully sparse occupancy network, termed SparseOcc. SparseOcc initially reconstructs a sparse 3D representation from camera-only inputs and subsequently predicts semantic/instance occupancy from the 3D sparse representation by sparse queries. A mask-guided sparse sampling is designed to enable sparse queries to interact with 2D features in a fully sparse manner, thereby circumventing costly dense features or global attention. Additionally, we design a thoughtful ray-based evaluation metric, namely RayIoU, to solve the inconsistency penalty along the depth axis raised in traditional voxel-level mIoU criteria. SparseOcc demonstrates its effectiveness by achieving a RayIoU of 34.0, while maintaining a real-time inference speed of 17.3 FPS, with 7 history frames inputs. By incorporating more preceding frames to 15, SparseOcc continuously improves its performance to 35.1 RayIoU without bells and whistles.
format Preprint
id arxiv_https___arxiv_org_abs_2312_17118
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Fully Sparse 3D Occupancy Prediction
Liu, Haisong
Chen, Yang
Wang, Haiguang
Yang, Zetong
Li, Tianyu
Zeng, Jia
Chen, Li
Li, Hongyang
Wang, Limin
Computer Vision and Pattern Recognition
Occupancy prediction plays a pivotal role in autonomous driving. Previous methods typically construct dense 3D volumes, neglecting the inherent sparsity of the scene and suffering from high computational costs. To bridge the gap, we introduce a novel fully sparse occupancy network, termed SparseOcc. SparseOcc initially reconstructs a sparse 3D representation from camera-only inputs and subsequently predicts semantic/instance occupancy from the 3D sparse representation by sparse queries. A mask-guided sparse sampling is designed to enable sparse queries to interact with 2D features in a fully sparse manner, thereby circumventing costly dense features or global attention. Additionally, we design a thoughtful ray-based evaluation metric, namely RayIoU, to solve the inconsistency penalty along the depth axis raised in traditional voxel-level mIoU criteria. SparseOcc demonstrates its effectiveness by achieving a RayIoU of 34.0, while maintaining a real-time inference speed of 17.3 FPS, with 7 history frames inputs. By incorporating more preceding frames to 15, SparseOcc continuously improves its performance to 35.1 RayIoU without bells and whistles.
title Fully Sparse 3D Occupancy Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.17118