MMLF: Multi-modal Multi-class Late Fusion for Object Detection with Uncertainty Estimation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Qihang, Zhao, Yang, Cheng, Hong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916433683808256
author Yang, Qihang
Zhao, Yang
Cheng, Hong
author_facet Yang, Qihang
Zhao, Yang
Cheng, Hong
contents Autonomous driving necessitates advanced object detection techniques that integrate information from multiple modalities to overcome the limitations associated with single-modal approaches. The challenges of aligning diverse data in early fusion and the complexities, along with overfitting issues introduced by deep fusion, underscore the efficacy of late fusion at the decision level. Late fusion ensures seamless integration without altering the original detector's network structure. This paper introduces a pioneering Multi-modal Multi-class Late Fusion method, designed for late fusion to enable multi-class detection. Fusion experiments conducted on the KITTI validation and official test datasets illustrate substantial performance improvements, presenting our model as a versatile solution for multi-modal object detection in autonomous driving. Moreover, our approach incorporates uncertainty analysis into the classification fusion process, rendering our model more transparent and trustworthy and providing more reliable insights into category predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08739
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMLF: Multi-modal Multi-class Late Fusion for Object Detection with Uncertainty Estimation
Yang, Qihang
Zhao, Yang
Cheng, Hong
Computer Vision and Pattern Recognition
Systems and Control
Autonomous driving necessitates advanced object detection techniques that integrate information from multiple modalities to overcome the limitations associated with single-modal approaches. The challenges of aligning diverse data in early fusion and the complexities, along with overfitting issues introduced by deep fusion, underscore the efficacy of late fusion at the decision level. Late fusion ensures seamless integration without altering the original detector's network structure. This paper introduces a pioneering Multi-modal Multi-class Late Fusion method, designed for late fusion to enable multi-class detection. Fusion experiments conducted on the KITTI validation and official test datasets illustrate substantial performance improvements, presenting our model as a versatile solution for multi-modal object detection in autonomous driving. Moreover, our approach incorporates uncertainty analysis into the classification fusion process, rendering our model more transparent and trustworthy and providing more reliable insights into category predictions.
title MMLF: Multi-modal Multi-class Late Fusion for Object Detection with Uncertainty Estimation
topic Computer Vision and Pattern Recognition
Systems and Control
url https://arxiv.org/abs/2410.08739