MapFM: Foundation Model-Driven HD Mapping with Multi-Task Contextual Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ivanov, Leonid, Yuryev, Vasily, Yudin, Dmitry
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912437967519744
author Ivanov, Leonid
Yuryev, Vasily
Yudin, Dmitry
author_facet Ivanov, Leonid
Yuryev, Vasily
Yudin, Dmitry
contents In autonomous driving, high-definition (HD) maps and semantic maps in bird's-eye view (BEV) are essential for accurate localization, planning, and decision-making. This paper introduces an enhanced End-to-End model named MapFM for online vectorized HD map generation. We show significantly boost feature representation quality by incorporating powerful foundation model for encoding camera images. To further enrich the model's understanding of the environment and improve prediction quality, we integrate auxiliary prediction heads for semantic segmentation in the BEV representation. This multi-task learning approach provides richer contextual supervision, leading to a more comprehensive scene representation and ultimately resulting in higher accuracy and improved quality of the predicted vectorized HD maps. The source code is available at https://github.com/LIvanoff/MapFM.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15313
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MapFM: Foundation Model-Driven HD Mapping with Multi-Task Contextual Learning
Ivanov, Leonid
Yuryev, Vasily
Yudin, Dmitry
Computer Vision and Pattern Recognition
Artificial Intelligence
In autonomous driving, high-definition (HD) maps and semantic maps in bird's-eye view (BEV) are essential for accurate localization, planning, and decision-making. This paper introduces an enhanced End-to-End model named MapFM for online vectorized HD map generation. We show significantly boost feature representation quality by incorporating powerful foundation model for encoding camera images. To further enrich the model's understanding of the environment and improve prediction quality, we integrate auxiliary prediction heads for semantic segmentation in the BEV representation. This multi-task learning approach provides richer contextual supervision, leading to a more comprehensive scene representation and ultimately resulting in higher accuracy and improved quality of the predicted vectorized HD maps. The source code is available at https://github.com/LIvanoff/MapFM.
title MapFM: Foundation Model-Driven HD Mapping with Multi-Task Contextual Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.15313