Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lai, Delun, Zhang, Yeyubei, Liu, Yunchong, Li, Chaojie, Mo, Huadong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913810085838848
author Lai, Delun
Zhang, Yeyubei
Liu, Yunchong
Li, Chaojie
Mo, Huadong
author_facet Lai, Delun
Zhang, Yeyubei
Liu, Yunchong
Li, Chaojie
Mo, Huadong
contents This paper introduces a novel deep learning-based multimodal fusion architecture aimed at enhancing the perception capabilities of autonomous navigation robots in complex environments. By utilizing innovative feature extraction modules, adaptive fusion strategies, and time-series modeling mechanisms, the system effectively integrates RGB images and LiDAR data. The key contributions of this work are as follows: a. the design of a lightweight feature extraction network to enhance feature representation; b. the development of an adaptive weighted cross-modal fusion strategy to improve system robustness; and c. the incorporation of time-series information modeling to boost dynamic scene perception accuracy. Experimental results on the KITTI dataset demonstrate that the proposed approach increases navigation and positioning accuracy by 3.5% and 2.2%, respectively, while maintaining real-time performance. This work provides a novel solution for autonomous robot navigation in complex environments.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19002
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation
Lai, Delun
Zhang, Yeyubei
Liu, Yunchong
Li, Chaojie
Mo, Huadong
Machine Learning
Computer Vision and Pattern Recognition
Robotics
This paper introduces a novel deep learning-based multimodal fusion architecture aimed at enhancing the perception capabilities of autonomous navigation robots in complex environments. By utilizing innovative feature extraction modules, adaptive fusion strategies, and time-series modeling mechanisms, the system effectively integrates RGB images and LiDAR data. The key contributions of this work are as follows: a. the design of a lightweight feature extraction network to enhance feature representation; b. the development of an adaptive weighted cross-modal fusion strategy to improve system robustness; and c. the incorporation of time-series information modeling to boost dynamic scene perception accuracy. Experimental results on the KITTI dataset demonstrate that the proposed approach increases navigation and positioning accuracy by 3.5% and 2.2%, respectively, while maintaining real-time performance. This work provides a novel solution for autonomous robot navigation in complex environments.
title Deep Learning-Based Multi-Modal Fusion for Robust Robot Perception and Navigation
topic Machine Learning
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2504.19002