OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zheng, Lianqing, Yang, Long, Lin, Qunshu, Ai, Wenjin, Liu, Minghao, Lu, Shouyi, Liu, Jianan, Ren, Hongze, Mo, Jingyue, Bai, Xiaokai, Bai, Jie, Ma, Zhixiong, Zhu, Xichan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912889126780928
author Zheng, Lianqing
Yang, Long
Lin, Qunshu
Ai, Wenjin
Liu, Minghao
Lu, Shouyi
Liu, Jianan
Ren, Hongze
Mo, Jingyue
Bai, Xiaokai
Bai, Jie
Ma, Zhixiong
Zhu, Xichan
author_facet Zheng, Lianqing
Yang, Long
Lin, Qunshu
Ai, Wenjin
Liu, Minghao
Lu, Shouyi
Liu, Jianan
Ren, Hongze
Mo, Jingyue
Bai, Xiaokai
Bai, Jie
Ma, Zhixiong
Zhu, Xichan
contents The rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the development of effective data-driven autonomous driving solutions. Next-generation autonomous driving datasets must be multimodal, incorporating data from advanced sensors that feature extensive data coverage, detailed annotations, and diverse scene representation. To address this need, we present OmniHD-Scenes, a large-scale multimodal dataset that provides comprehensive omnidirectional high-definition data. The OmniHD-Scenes dataset combines data from 128-beam LiDAR, six cameras, and six 4D imaging radar systems to achieve full environmental perception. The dataset comprises 1501 clips, each approximately 30-s long, totaling more than 450K synchronized frames and more than 5.85 million synchronized sensor data points. We also propose a novel 4D annotation pipeline. To date, we have annotated 200 clips with more than 514K precise 3D bounding boxes. These clips also include semantic segmentation annotations for static scene elements. Additionally, we introduce a novel automated pipeline for generation of the dense occupancy ground truth, which effectively leverages information from non-key frames. Alongside the proposed dataset, we establish comprehensive evaluation metrics, baseline models, and benchmarks for 3D detection and semantic occupancy prediction. These benchmarks utilize surround-view cameras and 4D imaging radar to explore cost-effective sensor solutions for autonomous driving applications. Extensive experiments demonstrate the effectiveness of our low-cost sensor configuration and its robustness under adverse conditions. Data will be released at https://www.2077ai.com/OmniHD-Scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10734
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
Zheng, Lianqing
Yang, Long
Lin, Qunshu
Ai, Wenjin
Liu, Minghao
Lu, Shouyi
Liu, Jianan
Ren, Hongze
Mo, Jingyue
Bai, Xiaokai
Bai, Jie
Ma, Zhixiong
Zhu, Xichan
Computer Vision and Pattern Recognition
The rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the development of effective data-driven autonomous driving solutions. Next-generation autonomous driving datasets must be multimodal, incorporating data from advanced sensors that feature extensive data coverage, detailed annotations, and diverse scene representation. To address this need, we present OmniHD-Scenes, a large-scale multimodal dataset that provides comprehensive omnidirectional high-definition data. The OmniHD-Scenes dataset combines data from 128-beam LiDAR, six cameras, and six 4D imaging radar systems to achieve full environmental perception. The dataset comprises 1501 clips, each approximately 30-s long, totaling more than 450K synchronized frames and more than 5.85 million synchronized sensor data points. We also propose a novel 4D annotation pipeline. To date, we have annotated 200 clips with more than 514K precise 3D bounding boxes. These clips also include semantic segmentation annotations for static scene elements. Additionally, we introduce a novel automated pipeline for generation of the dense occupancy ground truth, which effectively leverages information from non-key frames. Alongside the proposed dataset, we establish comprehensive evaluation metrics, baseline models, and benchmarks for 3D detection and semantic occupancy prediction. These benchmarks utilize surround-view cameras and 4D imaging radar to explore cost-effective sensor solutions for autonomous driving applications. Extensive experiments demonstrate the effectiveness of our low-cost sensor configuration and its robustness under adverse conditions. Data will be released at https://www.2077ai.com/OmniHD-Scenes.
title OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.10734