Open Vocabulary Monocular 3D Object Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yao, Jin, Gu, Hao, Chen, Xuweiyi, Wang, Jiayun, Cheng, Zezhou
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911286934110208
author Yao, Jin
Gu, Hao
Chen, Xuweiyi
Wang, Jiayun
Cheng, Zezhou
author_facet Yao, Jin
Gu, Hao
Chen, Xuweiyi
Wang, Jiayun
Cheng, Zezhou
contents We propose and study open-vocabulary monocular 3D detection, a novel task that aims to detect objects of any categores in metric 3D space from a single RGB image. Existing 3D object detectors either rely on costly sensors such as LiDAR or multi-view setups, or remain confined to closed vocabularies settings with limited categories, restricting their applicability. We identify two key challenges in this new setting. First, the scarcity of 3D bounding box annotations limits the ability to train generalizable models. To reduce dependence on 3D supervision, we propose a framework that effectively integrates pretrained 2D and 3D vision foundation models. Second, missing labels and semantic ambiguities (\eg, table vs. desk) in existing datasets hinder reliable evaluation. To address this, we design a novel metric that captures model performance while mitigating annotation issues. Our approach achieves state-of-the-art results in zero-shot 3D detection of novel categories as well as in-domain detection on seen classes. We hope our method provides a strong baseline and our evaluation protocol establishes a reliable benchmark for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16833
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Open Vocabulary Monocular 3D Object Detection
Yao, Jin
Gu, Hao
Chen, Xuweiyi
Wang, Jiayun
Cheng, Zezhou
Computer Vision and Pattern Recognition
We propose and study open-vocabulary monocular 3D detection, a novel task that aims to detect objects of any categores in metric 3D space from a single RGB image. Existing 3D object detectors either rely on costly sensors such as LiDAR or multi-view setups, or remain confined to closed vocabularies settings with limited categories, restricting their applicability. We identify two key challenges in this new setting. First, the scarcity of 3D bounding box annotations limits the ability to train generalizable models. To reduce dependence on 3D supervision, we propose a framework that effectively integrates pretrained 2D and 3D vision foundation models. Second, missing labels and semantic ambiguities (\eg, table vs. desk) in existing datasets hinder reliable evaluation. To address this, we design a novel metric that captures model performance while mitigating annotation issues. Our approach achieves state-of-the-art results in zero-shot 3D detection of novel categories as well as in-domain detection on seen classes. We hope our method provides a strong baseline and our evaluation protocol establishes a reliable benchmark for future research.
title Open Vocabulary Monocular 3D Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16833