M3TR: A Generalist Model for Real-World HD Map Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Immel, Fabian, Fehler, Richard, Bieder, Frank, Pauls, Jan-Hendrik, Stiller, Christoph
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910959355822080
author Immel, Fabian
Fehler, Richard
Bieder, Frank
Pauls, Jan-Hendrik
Stiller, Christoph
author_facet Immel, Fabian
Fehler, Richard
Bieder, Frank
Pauls, Jan-Hendrik
Stiller, Christoph
contents Autonomous vehicles rely on HD maps for their operation, but offline HD maps eventually become outdated. For this reason, online HD map construction methods use live sensor data to infer map information instead. Research on real map changes shows that oftentimes entire parts of an HD map remain unchanged and can be used as a prior. We therefore introduce M3TR (Multi-Masking Map Transformer), a generalist approach for HD map completion both with and without offline HD map priors. As a necessary foundation, we address shortcomings in ground truth labels for Argoverse 2 and nuScenes and propose the first comprehensive benchmark for HD map completion. Unlike existing models that specialize in a single kind of map change, which is unrealistic for deployment, our Generalist model handles all kinds of changes, matching the effectiveness of Expert models. With our map masking as augmentation regime, we can even achieve a +1.4 mAP improvement without a prior. Finally, by fully utilizing prior HD map elements and optimizing query designs, M3TR outperforms existing methods by +4.3 mAP while being the first real-world deployable model for offline HD map priors. Code is available at https://github.com/immel-f/m3tr
format Preprint
id arxiv_https___arxiv_org_abs_2411_10316
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle M3TR: A Generalist Model for Real-World HD Map Completion
Immel, Fabian
Fehler, Richard
Bieder, Frank
Pauls, Jan-Hendrik
Stiller, Christoph
Computer Vision and Pattern Recognition
Robotics
Autonomous vehicles rely on HD maps for their operation, but offline HD maps eventually become outdated. For this reason, online HD map construction methods use live sensor data to infer map information instead. Research on real map changes shows that oftentimes entire parts of an HD map remain unchanged and can be used as a prior. We therefore introduce M3TR (Multi-Masking Map Transformer), a generalist approach for HD map completion both with and without offline HD map priors. As a necessary foundation, we address shortcomings in ground truth labels for Argoverse 2 and nuScenes and propose the first comprehensive benchmark for HD map completion. Unlike existing models that specialize in a single kind of map change, which is unrealistic for deployment, our Generalist model handles all kinds of changes, matching the effectiveness of Expert models. With our map masking as augmentation regime, we can even achieve a +1.4 mAP improvement without a prior. Finally, by fully utilizing prior HD map elements and optimizing query designs, M3TR outperforms existing methods by +4.3 mAP while being the first real-world deployable model for offline HD map priors. Code is available at https://github.com/immel-f/m3tr
title M3TR: A Generalist Model for Real-World HD Map Completion
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2411.10316