Weak Cube R-CNN: Weakly Supervised 3D Detection using only 2D Bounding Boxes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hansen, Andreas Lau, Wanzeck, Lukas, Papadopoulos, Dim P.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912333674053632
author Hansen, Andreas Lau
Wanzeck, Lukas
Papadopoulos, Dim P.
author_facet Hansen, Andreas Lau
Wanzeck, Lukas
Papadopoulos, Dim P.
contents Monocular 3D object detection is an essential task in computer vision, and it has several applications in robotics and virtual reality. However, 3D object detectors are typically trained in a fully supervised way, relying extensively on 3D labeled data, which is labor-intensive and costly to annotate. This work focuses on weakly-supervised 3D detection to reduce data needs using a monocular method that leverages a singlecamera system over expensive LiDAR sensors or multi-camera setups. We propose a general model Weak Cube R-CNN, which can predict objects in 3D at inference time, requiring only 2D box annotations for training by exploiting the relationship between 2D projections of 3D cubes. Our proposed method utilizes pre-trained frozen foundation 2D models to estimate depth and orientation information on a training set. We use these estimated values as pseudo-ground truths during training. We design loss functions that avoid 3D labels by incorporating information from the external models into the loss. In this way, we aim to implicitly transfer knowledge from these large foundation 2D models without having access to 3D bounding box annotations. Experimental results on the SUN RGB-D dataset show increased performance in accuracy compared to an annotation time equalized Cube R-CNN baseline. While not precise for centimetre-level measurements, this method provides a strong foundation for further research.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13297
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Weak Cube R-CNN: Weakly Supervised 3D Detection using only 2D Bounding Boxes
Hansen, Andreas Lau
Wanzeck, Lukas
Papadopoulos, Dim P.
Computer Vision and Pattern Recognition
I.4
Monocular 3D object detection is an essential task in computer vision, and it has several applications in robotics and virtual reality. However, 3D object detectors are typically trained in a fully supervised way, relying extensively on 3D labeled data, which is labor-intensive and costly to annotate. This work focuses on weakly-supervised 3D detection to reduce data needs using a monocular method that leverages a singlecamera system over expensive LiDAR sensors or multi-camera setups. We propose a general model Weak Cube R-CNN, which can predict objects in 3D at inference time, requiring only 2D box annotations for training by exploiting the relationship between 2D projections of 3D cubes. Our proposed method utilizes pre-trained frozen foundation 2D models to estimate depth and orientation information on a training set. We use these estimated values as pseudo-ground truths during training. We design loss functions that avoid 3D labels by incorporating information from the external models into the loss. In this way, we aim to implicitly transfer knowledge from these large foundation 2D models without having access to 3D bounding box annotations. Experimental results on the SUN RGB-D dataset show increased performance in accuracy compared to an annotation time equalized Cube R-CNN baseline. While not precise for centimetre-level measurements, this method provides a strong foundation for further research.
title Weak Cube R-CNN: Weakly Supervised 3D Detection using only 2D Bounding Boxes
topic Computer Vision and Pattern Recognition
I.4
url https://arxiv.org/abs/2504.13297