TopoBDA: Towards Bezier Deformable Attention for Road Topology Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kalfaoglu, Muhammet Esat, Ozturk, Halil Ibrahim, Kilinc, Ozsel, Temizel, Alptekin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917188648042496
author Kalfaoglu, Muhammet Esat
Ozturk, Halil Ibrahim
Kilinc, Ozsel
Temizel, Alptekin
author_facet Kalfaoglu, Muhammet Esat
Ozturk, Halil Ibrahim
Kilinc, Ozsel
Temizel, Alptekin
contents Understanding road topology is crucial for autonomous driving. This paper introduces TopoBDA (Topology with Bezier Deformable Attention), a novel approach that enhances road topology comprehension by leveraging Bezier Deformable Attention (BDA). TopoBDA processes multi-camera 360-degree imagery to generate Bird's Eye View (BEV) features, which are refined through a transformer decoder employing BDA. BDA utilizes Bezier control points to drive the deformable attention mechanism, improving the detection and representation of elongated and thin polyline structures, such as lane centerlines. Additionally, TopoBDA integrates two auxiliary components: an instance mask formulation loss and a one-to-many set prediction loss strategy, to further refine centerline detection and enhance road topology understanding. Experimental evaluations on the OpenLane-V2 dataset demonstrate that TopoBDA outperforms existing methods, achieving state-of-the-art results in centerline detection and topology reasoning. TopoBDA also achieves the best results on the OpenLane-V1 dataset in 3D lane detection. Further experiments on integrating multi-modal data -- such as LiDAR, radar, and SDMap -- show that multimodal inputs can further enhance performance in road topology understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18951
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TopoBDA: Towards Bezier Deformable Attention for Road Topology Understanding
Kalfaoglu, Muhammet Esat
Ozturk, Halil Ibrahim
Kilinc, Ozsel
Temizel, Alptekin
Computer Vision and Pattern Recognition
Understanding road topology is crucial for autonomous driving. This paper introduces TopoBDA (Topology with Bezier Deformable Attention), a novel approach that enhances road topology comprehension by leveraging Bezier Deformable Attention (BDA). TopoBDA processes multi-camera 360-degree imagery to generate Bird's Eye View (BEV) features, which are refined through a transformer decoder employing BDA. BDA utilizes Bezier control points to drive the deformable attention mechanism, improving the detection and representation of elongated and thin polyline structures, such as lane centerlines. Additionally, TopoBDA integrates two auxiliary components: an instance mask formulation loss and a one-to-many set prediction loss strategy, to further refine centerline detection and enhance road topology understanding. Experimental evaluations on the OpenLane-V2 dataset demonstrate that TopoBDA outperforms existing methods, achieving state-of-the-art results in centerline detection and topology reasoning. TopoBDA also achieves the best results on the OpenLane-V1 dataset in 3D lane detection. Further experiments on integrating multi-modal data -- such as LiDAR, radar, and SDMap -- show that multimodal inputs can further enhance performance in road topology understanding.
title TopoBDA: Towards Bezier Deformable Attention for Road Topology Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.18951