Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Han, Xiaofeng, Chen, Shunpeng, Fu, Zenghuang, Feng, Zhe, Fan, Lue, An, Dong, Wang, Changwei, Guo, Li, Meng, Weiliang, Zhang, Xiaopeng, Xu, Rongtao, Xu, Shibiao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914094484815872
author Han, Xiaofeng
Chen, Shunpeng
Fu, Zenghuang
Feng, Zhe
Fan, Lue
An, Dong
Wang, Changwei
Guo, Li
Meng, Weiliang
Zhang, Xiaopeng
Xu, Rongtao
Xu, Shibiao
author_facet Han, Xiaofeng
Chen, Shunpeng
Fu, Zenghuang
Feng, Zhe
Fan, Lue
An, Dong
Wang, Changwei
Guo, Li
Meng, Weiliang
Zhang, Xiaopeng
Xu, Rongtao
Xu, Shibiao
contents Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion methods and VLMs in the field of robot vision. For semantic scene understanding tasks, we categorize fusion approaches into encoder-decoder frameworks, attention-based architectures, and graph neural networks. Meanwhile, we also analyze the architectural characteristics and practical implementations of these fusion strategies in key tasks such as simultaneous localization and mapping (SLAM), 3D object detection, navigation, and manipulation. We compare the evolutionary paths and applicability of VLMs based on large language models (LLMs) with traditional multimodal fusion methods.Additionally, we conduct an in-depth analysis of commonly used datasets, evaluating their applicability and challenges in real-world robotic scenarios. Building on this analysis, we identify key challenges in current research, including cross-modal alignment, efficient fusion, real-time deployment, and domain adaptation. We propose future directions such as self-supervised learning for robust multimodal representations, structured spatial memory and environment modeling to enhance spatial intelligence, and the integration of adversarial robustness and human feedback mechanisms to enable ethically aligned system deployment. Through a comprehensive review, comparative analysis, and forward-looking discussion, we provide a valuable reference for advancing multimodal perception and interaction in robotic vision. A comprehensive list of studies in this survey is available at https://github.com/Xiaofeng-Han-Res/MF-RV.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
Han, Xiaofeng
Chen, Shunpeng
Fu, Zenghuang
Feng, Zhe
Fan, Lue
An, Dong
Wang, Changwei
Guo, Li
Meng, Weiliang
Zhang, Xiaopeng
Xu, Rongtao
Xu, Shibiao
Robotics
Computer Vision and Pattern Recognition
Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion methods and VLMs in the field of robot vision. For semantic scene understanding tasks, we categorize fusion approaches into encoder-decoder frameworks, attention-based architectures, and graph neural networks. Meanwhile, we also analyze the architectural characteristics and practical implementations of these fusion strategies in key tasks such as simultaneous localization and mapping (SLAM), 3D object detection, navigation, and manipulation. We compare the evolutionary paths and applicability of VLMs based on large language models (LLMs) with traditional multimodal fusion methods.Additionally, we conduct an in-depth analysis of commonly used datasets, evaluating their applicability and challenges in real-world robotic scenarios. Building on this analysis, we identify key challenges in current research, including cross-modal alignment, efficient fusion, real-time deployment, and domain adaptation. We propose future directions such as self-supervised learning for robust multimodal representations, structured spatial memory and environment modeling to enhance spatial intelligence, and the integration of adversarial robustness and human feedback mechanisms to enable ethically aligned system deployment. Through a comprehensive review, comparative analysis, and forward-looking discussion, we provide a valuable reference for advancing multimodal perception and interaction in robotic vision. A comprehensive list of studies in this survey is available at https://github.com/Xiaofeng-Han-Res/MF-RV.
title Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.02477