CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shaw, Ankit Kumar, Jiang, Kun, Wen, Tuopu, Sah, Chandan Kumar, Shi, Yining, Yang, Mengmeng, Yang, Diange, Lian, Xiaoli
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910911737888768
author Shaw, Ankit Kumar
Jiang, Kun
Wen, Tuopu
Sah, Chandan Kumar
Shi, Yining
Yang, Mengmeng
Yang, Diange
Lian, Xiaoli
author_facet Shaw, Ankit Kumar
Jiang, Kun
Wen, Tuopu
Sah, Chandan Kumar
Shi, Yining
Yang, Mengmeng
Yang, Diange
Lian, Xiaoli
contents The rapid growth of intelligent connected vehicles (ICVs) and integrated vehicle-road-cloud systems has increased the demand for accurate, real-time HD map updates. However, ensuring map reliability remains challenging due to inconsistencies in crowdsourced data, which suffer from motion blur, lighting variations, adverse weather, and lane marking degradation. This paper introduces CleanMAP, a Multimodal Large Language Model (MLLM)-based distillation framework designed to filter and refine crowdsourced data for high-confidence HD map updates. CleanMAP leverages an MLLM-driven lane visibility scoring model that systematically quantifies key visual parameters, assigning confidence scores (0-10) based on their impact on lane detection. A novel dynamic piecewise confidence-scoring function adapts scores based on lane visibility, ensuring strong alignment with human evaluations while effectively filtering unreliable data. To further optimize map accuracy, a confidence-driven local map fusion strategy ranks and selects the top-k highest-scoring local maps within an optimal confidence range (best score minus 10%), striking a balance between data quality and quantity. Experimental evaluations on a real-world autonomous vehicle dataset validate CleanMAP's effectiveness, demonstrating that fusing the top three local maps achieves the lowest mean map update error of 0.28m, outperforming the baseline (0.37m) and meeting stringent accuracy thresholds (<= 0.32m). Further validation with real-vehicle data confirms 84.88% alignment with human evaluators, reinforcing the model's robustness and reliability. This work establishes CleanMAP as a scalable and deployable solution for crowdsourced HD map updates, ensuring more precise and reliable autonomous navigation. The code will be available at https://Ankit-Zefan.github.io/CleanMap/
format Preprint
id arxiv_https___arxiv_org_abs_2504_10738
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates
Shaw, Ankit Kumar
Jiang, Kun
Wen, Tuopu
Sah, Chandan Kumar
Shi, Yining
Yang, Mengmeng
Yang, Diange
Lian, Xiaoli
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
I.2.9; I.2.7; I.2.10; I.5.5; I.5.4; I.2.11
The rapid growth of intelligent connected vehicles (ICVs) and integrated vehicle-road-cloud systems has increased the demand for accurate, real-time HD map updates. However, ensuring map reliability remains challenging due to inconsistencies in crowdsourced data, which suffer from motion blur, lighting variations, adverse weather, and lane marking degradation. This paper introduces CleanMAP, a Multimodal Large Language Model (MLLM)-based distillation framework designed to filter and refine crowdsourced data for high-confidence HD map updates. CleanMAP leverages an MLLM-driven lane visibility scoring model that systematically quantifies key visual parameters, assigning confidence scores (0-10) based on their impact on lane detection. A novel dynamic piecewise confidence-scoring function adapts scores based on lane visibility, ensuring strong alignment with human evaluations while effectively filtering unreliable data. To further optimize map accuracy, a confidence-driven local map fusion strategy ranks and selects the top-k highest-scoring local maps within an optimal confidence range (best score minus 10%), striking a balance between data quality and quantity. Experimental evaluations on a real-world autonomous vehicle dataset validate CleanMAP's effectiveness, demonstrating that fusing the top three local maps achieves the lowest mean map update error of 0.28m, outperforming the baseline (0.37m) and meeting stringent accuracy thresholds (<= 0.32m). Further validation with real-vehicle data confirms 84.88% alignment with human evaluators, reinforcing the model's robustness and reliability. This work establishes CleanMAP as a scalable and deployable solution for crowdsourced HD map updates, ensuring more precise and reliable autonomous navigation. The code will be available at https://Ankit-Zefan.github.io/CleanMap/
title CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
I.2.9; I.2.7; I.2.10; I.5.5; I.5.4; I.2.11
url https://arxiv.org/abs/2504.10738