Multimodal HD Mapping for Intersections by Intelligent Roadside Units

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Zhongzhang, Fan, Miao, Xu, Shengtong, Yang, Mengmeng, Jiang, Kun, Liu, Xiangzeng, Xiong, Haoyi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915385591201792
author Chen, Zhongzhang
Fan, Miao
Xu, Shengtong
Yang, Mengmeng
Jiang, Kun
Liu, Xiangzeng
Xiong, Haoyi
author_facet Chen, Zhongzhang
Fan, Miao
Xu, Shengtong
Yang, Mengmeng
Jiang, Kun
Liu, Xiangzeng
Xiong, Haoyi
contents High-definition (HD) semantic mapping of complex intersections poses significant challenges for traditional vehicle-based approaches due to occlusions and limited perspectives. This paper introduces a novel camera-LiDAR fusion framework that leverages elevated intelligent roadside units (IRUs). Additionally, we present RS-seq, a comprehensive dataset developed through the systematic enhancement and annotation of the V2X-Seq dataset. RS-seq includes precisely labelled camera imagery and LiDAR point clouds collected from roadside installations, along with vectorized maps for seven intersections annotated with detailed features such as lane dividers, pedestrian crossings, and stop lines. This dataset facilitates the systematic investigation of cross-modal complementarity for HD map generation using IRU data. The proposed fusion framework employs a two-stage process that integrates modality-specific feature extraction and cross-modal semantic integration, capitalizing on camera high-resolution texture and precise geometric data from LiDAR. Quantitative evaluations using the RS-seq dataset demonstrate that our multimodal approach consistently surpasses unimodal methods. Specifically, compared to unimodal baselines evaluated on the RS-seq dataset, the multimodal approach improves the mean Intersection-over-Union (mIoU) for semantic segmentation by 4\% over the image-only results and 18\% over the point cloud-only results. This study establishes a baseline methodology for IRU-based HD semantic mapping and provides a valuable dataset for future research in infrastructure-assisted autonomous driving systems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08903
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal HD Mapping for Intersections by Intelligent Roadside Units
Chen, Zhongzhang
Fan, Miao
Xu, Shengtong
Yang, Mengmeng
Jiang, Kun
Liu, Xiangzeng
Xiong, Haoyi
Robotics
Computer Vision and Pattern Recognition
High-definition (HD) semantic mapping of complex intersections poses significant challenges for traditional vehicle-based approaches due to occlusions and limited perspectives. This paper introduces a novel camera-LiDAR fusion framework that leverages elevated intelligent roadside units (IRUs). Additionally, we present RS-seq, a comprehensive dataset developed through the systematic enhancement and annotation of the V2X-Seq dataset. RS-seq includes precisely labelled camera imagery and LiDAR point clouds collected from roadside installations, along with vectorized maps for seven intersections annotated with detailed features such as lane dividers, pedestrian crossings, and stop lines. This dataset facilitates the systematic investigation of cross-modal complementarity for HD map generation using IRU data. The proposed fusion framework employs a two-stage process that integrates modality-specific feature extraction and cross-modal semantic integration, capitalizing on camera high-resolution texture and precise geometric data from LiDAR. Quantitative evaluations using the RS-seq dataset demonstrate that our multimodal approach consistently surpasses unimodal methods. Specifically, compared to unimodal baselines evaluated on the RS-seq dataset, the multimodal approach improves the mean Intersection-over-Union (mIoU) for semantic segmentation by 4\% over the image-only results and 18\% over the point cloud-only results. This study establishes a baseline methodology for IRU-based HD semantic mapping and provides a valuable dataset for future research in infrastructure-assisted autonomous driving systems.
title Multimodal HD Mapping for Intersections by Intelligent Roadside Units
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.08903