MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hou, Minghui, Huang, Wei-Hsing, Liang, Shaofeng, Liu, Daizong, Wen, Tai-Hao, Wang, Gang, Guan, Runwei, Ding, Weiping
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909964522487808
author Hou, Minghui
Huang, Wei-Hsing
Liang, Shaofeng
Liu, Daizong
Wen, Tai-Hao
Wang, Gang
Guan, Runwei
Ding, Weiping
author_facet Hou, Minghui
Huang, Wei-Hsing
Liang, Shaofeng
Liu, Daizong
Wen, Tai-Hao
Wang, Gang
Guan, Runwei
Ding, Weiping
contents Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are constrained by the image understanding paradigm in 2D plane, which restricts their capability to perceive 3D spatial information and perform deep semantic fusion, resulting in suboptimal performance in complex autonomous driving environments. This study proposes MMDrive, an multimodal vision-language model framework that extends traditional image understanding to a generalized 3D scene understanding framework. MMDrive incorporates three complementary modalities, including occupancy maps, LiDAR point clouds, and textual scene descriptions. To this end, it introduces two novel components for adaptive cross-modal fusion and key information extraction. Specifically, the Text-oriented Multimodal Modulator dynamically weights the contributions of each modality based on the semantic cues in the question, guiding context-aware feature integration. The Cross-Modal Abstractor employs learnable abstract tokens to generate compact, cross-modal summaries that highlight key regions and essential semantics. Comprehensive evaluations on the DriveLM and NuScenes-QA benchmarks demonstrate that MMDrive achieves significant performance gains over existing vision-language models for autonomous driving, with a BLEU-4 score of 54.56 and METEOR of 41.78 on DriveLM, and an accuracy score of 62.7% on NuScenes-QA. MMDrive effectively breaks the traditional image-only understanding barrier, enabling robust multimodal reasoning in complex driving environments and providing a new foundation for interpretable autonomous driving scene understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13177
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
Hou, Minghui
Huang, Wei-Hsing
Liang, Shaofeng
Liu, Daizong
Wen, Tai-Hao
Wang, Gang
Guan, Runwei
Ding, Weiping
Computer Vision and Pattern Recognition
Robotics
Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-language models are constrained by the image understanding paradigm in 2D plane, which restricts their capability to perceive 3D spatial information and perform deep semantic fusion, resulting in suboptimal performance in complex autonomous driving environments. This study proposes MMDrive, an multimodal vision-language model framework that extends traditional image understanding to a generalized 3D scene understanding framework. MMDrive incorporates three complementary modalities, including occupancy maps, LiDAR point clouds, and textual scene descriptions. To this end, it introduces two novel components for adaptive cross-modal fusion and key information extraction. Specifically, the Text-oriented Multimodal Modulator dynamically weights the contributions of each modality based on the semantic cues in the question, guiding context-aware feature integration. The Cross-Modal Abstractor employs learnable abstract tokens to generate compact, cross-modal summaries that highlight key regions and essential semantics. Comprehensive evaluations on the DriveLM and NuScenes-QA benchmarks demonstrate that MMDrive achieves significant performance gains over existing vision-language models for autonomous driving, with a BLEU-4 score of 54.56 and METEOR of 41.78 on DriveLM, and an accuracy score of 62.7% on NuScenes-QA. MMDrive effectively breaks the traditional image-only understanding barrier, enabling robust multimodal reasoning in complex driving environments and providing a new foundation for interpretable autonomous driving scene understanding.
title MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.13177