Better Monocular 3D Detectors with LiDAR from the Past

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: You, Yurong, Phoo, Cheng Perng, Diaz-Ruiz, Carlos Andres, Luo, Katie Z, Chao, Wei-Lun, Campbell, Mark, Hariharan, Bharath, Weinberger, Kilian Q
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911833896517632
author You, Yurong
Phoo, Cheng Perng
Diaz-Ruiz, Carlos Andres
Luo, Katie Z
Chao, Wei-Lun
Campbell, Mark
Hariharan, Bharath
Weinberger, Kilian Q
author_facet You, Yurong
Phoo, Cheng Perng
Diaz-Ruiz, Carlos Andres
Luo, Katie Z
Chao, Wei-Lun
Campbell, Mark
Hariharan, Bharath
Weinberger, Kilian Q
contents Accurate 3D object detection is crucial to autonomous driving. Though LiDAR-based detectors have achieved impressive performance, the high cost of LiDAR sensors precludes their widespread adoption in affordable vehicles. Camera-based detectors are cheaper alternatives but often suffer inferior performance compared to their LiDAR-based counterparts due to inherent depth ambiguities in images. In this work, we seek to improve monocular 3D detectors by leveraging unlabeled historical LiDAR data. Specifically, at inference time, we assume that the camera-based detectors have access to multiple unlabeled LiDAR scans from past traversals at locations of interest (potentially from other high-end vehicles equipped with LiDAR sensors). Under this setup, we proposed a novel, simple, and end-to-end trainable framework, termed AsyncDepth, to effectively extract relevant features from asynchronous LiDAR traversals of the same location for monocular 3D detectors. We show consistent and significant performance gain (up to 9 AP) across multiple state-of-the-art models and datasets with a negligible additional latency of 9.66 ms and a small storage cost.
format Preprint
id arxiv_https___arxiv_org_abs_2404_05139
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Better Monocular 3D Detectors with LiDAR from the Past
You, Yurong
Phoo, Cheng Perng
Diaz-Ruiz, Carlos Andres
Luo, Katie Z
Chao, Wei-Lun
Campbell, Mark
Hariharan, Bharath
Weinberger, Kilian Q
Computer Vision and Pattern Recognition
Robotics
Accurate 3D object detection is crucial to autonomous driving. Though LiDAR-based detectors have achieved impressive performance, the high cost of LiDAR sensors precludes their widespread adoption in affordable vehicles. Camera-based detectors are cheaper alternatives but often suffer inferior performance compared to their LiDAR-based counterparts due to inherent depth ambiguities in images. In this work, we seek to improve monocular 3D detectors by leveraging unlabeled historical LiDAR data. Specifically, at inference time, we assume that the camera-based detectors have access to multiple unlabeled LiDAR scans from past traversals at locations of interest (potentially from other high-end vehicles equipped with LiDAR sensors). Under this setup, we proposed a novel, simple, and end-to-end trainable framework, termed AsyncDepth, to effectively extract relevant features from asynchronous LiDAR traversals of the same location for monocular 3D detectors. We show consistent and significant performance gain (up to 9 AP) across multiple state-of-the-art models and datasets with a negligible additional latency of 9.66 ms and a small storage cost.
title Better Monocular 3D Detectors with LiDAR from the Past
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2404.05139