BEVMOSNet: Multimodal Fusion for BEV Moving Object Segmentation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cong, Hiep Truong, Sigatapu, Ajay Kumar, Das, Arindam, Sharma, Yashwanth, Satagopan, Venkatesh, Sistu, Ganesh, Eising, Ciaran
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912260468768768
author Cong, Hiep Truong
Sigatapu, Ajay Kumar
Das, Arindam
Sharma, Yashwanth
Satagopan, Venkatesh
Sistu, Ganesh
Eising, Ciaran
author_facet Cong, Hiep Truong
Sigatapu, Ajay Kumar
Das, Arindam
Sharma, Yashwanth
Satagopan, Venkatesh
Sistu, Ganesh
Eising, Ciaran
contents Accurate motion understanding of the dynamic objects within the scene in bird's-eye-view (BEV) is critical to ensure a reliable obstacle avoidance system and smooth path planning for autonomous vehicles. However, this task has received relatively limited exploration when compared to object detection and segmentation with only a few recent vision-based approaches presenting preliminary findings that significantly deteriorate in low-light, nighttime, and adverse weather conditions such as rain. Conversely, LiDAR and radar sensors remain almost unaffected in these scenarios, and radar provides key velocity information of the objects. Therefore, we introduce BEVMOSNet, to our knowledge, the first end-to-end multimodal fusion leveraging cameras, LiDAR, and radar to precisely predict the moving objects in BEV. In addition, we perform a deeper analysis to find out the optimal strategy for deformable cross-attention-guided sensor fusion for cross-sensor knowledge sharing in BEV. While evaluating BEVMOSNet on the nuScenes dataset, we show an overall improvement in IoU score of 36.59% compared to the vision-based unimodal baseline BEV-MoSeg (Sigatapu et al., 2023), and 2.35% compared to the multimodel SimpleBEV (Harley et al., 2022), extended for the motion segmentation task, establishing this method as the state-of-the-art in BEV motion segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03280
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BEVMOSNet: Multimodal Fusion for BEV Moving Object Segmentation
Cong, Hiep Truong
Sigatapu, Ajay Kumar
Das, Arindam
Sharma, Yashwanth
Satagopan, Venkatesh
Sistu, Ganesh
Eising, Ciaran
Computer Vision and Pattern Recognition
Accurate motion understanding of the dynamic objects within the scene in bird's-eye-view (BEV) is critical to ensure a reliable obstacle avoidance system and smooth path planning for autonomous vehicles. However, this task has received relatively limited exploration when compared to object detection and segmentation with only a few recent vision-based approaches presenting preliminary findings that significantly deteriorate in low-light, nighttime, and adverse weather conditions such as rain. Conversely, LiDAR and radar sensors remain almost unaffected in these scenarios, and radar provides key velocity information of the objects. Therefore, we introduce BEVMOSNet, to our knowledge, the first end-to-end multimodal fusion leveraging cameras, LiDAR, and radar to precisely predict the moving objects in BEV. In addition, we perform a deeper analysis to find out the optimal strategy for deformable cross-attention-guided sensor fusion for cross-sensor knowledge sharing in BEV. While evaluating BEVMOSNet on the nuScenes dataset, we show an overall improvement in IoU score of 36.59% compared to the vision-based unimodal baseline BEV-MoSeg (Sigatapu et al., 2023), and 2.35% compared to the multimodel SimpleBEV (Harley et al., 2022), extended for the motion segmentation task, establishing this method as the state-of-the-art in BEV motion segmentation.
title BEVMOSNet: Multimodal Fusion for BEV Moving Object Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.03280