AYDIV: Adaptable Yielding 3D Object Detection via Integrated Contextual Vision Transformer

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dam, Tanmoy, Dharavath, Sanjay Bhargav, Alam, Sameer, Lilith, Nimrod, Chakraborty, Supriyo, Feroskhan, Mir
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929480209006592
author Dam, Tanmoy
Dharavath, Sanjay Bhargav
Alam, Sameer
Lilith, Nimrod
Chakraborty, Supriyo
Feroskhan, Mir
author_facet Dam, Tanmoy
Dharavath, Sanjay Bhargav
Alam, Sameer
Lilith, Nimrod
Chakraborty, Supriyo
Feroskhan, Mir
contents Combining LiDAR and camera data has shown potential in enhancing short-distance object detection in autonomous driving systems. Yet, the fusion encounters difficulties with extended distance detection due to the contrast between LiDAR's sparse data and the dense resolution of cameras. Besides, discrepancies in the two data representations further complicate fusion methods. We introduce AYDIV, a novel framework integrating a tri-phase alignment process specifically designed to enhance long-distance detection even amidst data discrepancies. AYDIV consists of the Global Contextual Fusion Alignment Transformer (GCFAT), which improves the extraction of camera features and provides a deeper understanding of large-scale patterns; the Sparse Fused Feature Attention (SFFA), which fine-tunes the fusion of LiDAR and camera details; and the Volumetric Grid Attention (VGA) for a comprehensive spatial data fusion. AYDIV's performance on the Waymo Open Dataset (WOD) with an improvement of 1.24% in mAPH value(L2 difficulty) and the Argoverse2 Dataset with a performance improvement of 7.40% in AP value demonstrates its efficacy in comparison to other existing fusion-based methods. Our code is publicly available at https://github.com/sanjay-810/AYDIV2
format Preprint
id arxiv_https___arxiv_org_abs_2402_07680
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AYDIV: Adaptable Yielding 3D Object Detection via Integrated Contextual Vision Transformer
Dam, Tanmoy
Dharavath, Sanjay Bhargav
Alam, Sameer
Lilith, Nimrod
Chakraborty, Supriyo
Feroskhan, Mir
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Combining LiDAR and camera data has shown potential in enhancing short-distance object detection in autonomous driving systems. Yet, the fusion encounters difficulties with extended distance detection due to the contrast between LiDAR's sparse data and the dense resolution of cameras. Besides, discrepancies in the two data representations further complicate fusion methods. We introduce AYDIV, a novel framework integrating a tri-phase alignment process specifically designed to enhance long-distance detection even amidst data discrepancies. AYDIV consists of the Global Contextual Fusion Alignment Transformer (GCFAT), which improves the extraction of camera features and provides a deeper understanding of large-scale patterns; the Sparse Fused Feature Attention (SFFA), which fine-tunes the fusion of LiDAR and camera details; and the Volumetric Grid Attention (VGA) for a comprehensive spatial data fusion. AYDIV's performance on the Waymo Open Dataset (WOD) with an improvement of 1.24% in mAPH value(L2 difficulty) and the Argoverse2 Dataset with a performance improvement of 7.40% in AP value demonstrates its efficacy in comparison to other existing fusion-based methods. Our code is publicly available at https://github.com/sanjay-810/AYDIV2
title AYDIV: Adaptable Yielding 3D Object Detection via Integrated Contextual Vision Transformer
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2402.07680