MOSE: Boosting Vision-based Roadside 3D Object Detection with Scene Cues

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Xiahan, Chen, Mingjian, Tang, Sanli, Niu, Yi, Zhu, Jiang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916197579096064
author Chen, Xiahan
Chen, Mingjian
Tang, Sanli
Niu, Yi
Zhu, Jiang
author_facet Chen, Xiahan
Chen, Mingjian
Tang, Sanli
Niu, Yi
Zhu, Jiang
contents 3D object detection based on roadside cameras is an additional way for autonomous driving to alleviate the challenges of occlusion and short perception range from vehicle cameras. Previous methods for roadside 3D object detection mainly focus on modeling the depth or height of objects, neglecting the stationary of cameras and the characteristic of inter-frame consistency. In this work, we propose a novel framework, namely MOSE, for MOnocular 3D object detection with Scene cuEs. The scene cues are the frame-invariant scene-specific features, which are crucial for object localization and can be intuitively regarded as the height between the surface of the real road and the virtual ground plane. In the proposed framework, a scene cue bank is designed to aggregate scene cues from multiple frames of the same scene with a carefully designed extrinsic augmentation strategy. Then, a transformer-based decoder lifts the aggregated scene cues as well as the 3D position embeddings for 3D object location, which boosts generalization ability in heterologous scenes. The extensive experiment results on two public benchmarks demonstrate the state-of-the-art performance of the proposed method, which surpasses the existing methods by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2404_05280
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MOSE: Boosting Vision-based Roadside 3D Object Detection with Scene Cues
Chen, Xiahan
Chen, Mingjian
Tang, Sanli
Niu, Yi
Zhu, Jiang
Computer Vision and Pattern Recognition
3D object detection based on roadside cameras is an additional way for autonomous driving to alleviate the challenges of occlusion and short perception range from vehicle cameras. Previous methods for roadside 3D object detection mainly focus on modeling the depth or height of objects, neglecting the stationary of cameras and the characteristic of inter-frame consistency. In this work, we propose a novel framework, namely MOSE, for MOnocular 3D object detection with Scene cuEs. The scene cues are the frame-invariant scene-specific features, which are crucial for object localization and can be intuitively regarded as the height between the surface of the real road and the virtual ground plane. In the proposed framework, a scene cue bank is designed to aggregate scene cues from multiple frames of the same scene with a carefully designed extrinsic augmentation strategy. Then, a transformer-based decoder lifts the aggregated scene cues as well as the 3D position embeddings for 3D object location, which boosts generalization ability in heterologous scenes. The extensive experiment results on two public benchmarks demonstrate the state-of-the-art performance of the proposed method, which surpasses the existing methods by a large margin.
title MOSE: Boosting Vision-based Roadside 3D Object Detection with Scene Cues
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.05280