IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data Domain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhe, Huo, Xiaoliang, Fan, Siqi, Liu, Jingjing, Zhang, Ya-Qin, Wang, Yan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913671041515520
author Wang, Zhe
Huo, Xiaoliang
Fan, Siqi
Liu, Jingjing
Zhang, Ya-Qin
Wang, Yan
author_facet Wang, Zhe
Huo, Xiaoliang
Fan, Siqi
Liu, Jingjing
Zhang, Ya-Qin
Wang, Yan
contents In autonomous driving, The perception capabilities of the ego-vehicle can be improved with roadside sensors, which can provide a holistic view of the environment. However, existing monocular detection methods designed for vehicle cameras are not suitable for roadside cameras due to viewpoint domain gaps. To bridge this gap and Improve ROAdside Monocular 3D object detection, we propose IROAM, a semantic-geometry decoupled contrastive learning framework, which takes vehicle-side and roadside data as input simultaneously. IROAM has two significant modules. In-Domain Query Interaction module utilizes a transformer to learn content and depth information for each domain and outputs object queries. Cross-Domain Query Enhancement To learn better feature representations from two domains, Cross-Domain Query Enhancement decouples queries into semantic and geometry parts and only the former is used for contrastive learning. Experiments demonstrate the effectiveness of IROAM in improving roadside detector's performance. The results validate that IROAM has the capabilities to learn cross-domain information.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18162
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data Domain
Wang, Zhe
Huo, Xiaoliang
Fan, Siqi
Liu, Jingjing
Zhang, Ya-Qin
Wang, Yan
Computer Vision and Pattern Recognition
Robotics
In autonomous driving, The perception capabilities of the ego-vehicle can be improved with roadside sensors, which can provide a holistic view of the environment. However, existing monocular detection methods designed for vehicle cameras are not suitable for roadside cameras due to viewpoint domain gaps. To bridge this gap and Improve ROAdside Monocular 3D object detection, we propose IROAM, a semantic-geometry decoupled contrastive learning framework, which takes vehicle-side and roadside data as input simultaneously. IROAM has two significant modules. In-Domain Query Interaction module utilizes a transformer to learn content and depth information for each domain and outputs object queries. Cross-Domain Query Enhancement To learn better feature representations from two domains, Cross-Domain Query Enhancement decouples queries into semantic and geometry parts and only the former is used for contrastive learning. Experiments demonstrate the effectiveness of IROAM in improving roadside detector's performance. The results validate that IROAM has the capabilities to learn cross-domain information.
title IROAM: Improving Roadside Monocular 3D Object Detection Learning from Autonomous Vehicle Data Domain
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2501.18162