Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Zongchuang, Fu, Haoyu, Liang, Dingkang, Zhou, Xin, Zhang, Dingyuan, Xie, Hongwei, Wang, Bing, Bai, Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912373878554624
author Zhao, Zongchuang
Fu, Haoyu
Liang, Dingkang
Zhou, Xin
Zhang, Dingyuan
Xie, Hongwei
Wang, Bing
Bai, Xiang
author_facet Zhao, Zongchuang
Fu, Haoyu
Liang, Dingkang
Zhou, Xin
Zhang, Dingyuan
Xie, Hongwei
Wang, Bing
Bai, Xiang
contents The Large Visual-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing research typically focuses on front-view perspectives and partial objects within scenes, struggling to achieve comprehensive scene understanding. Meanwhile, existing LVLMs suffer from the lack of mapping relationship between 2D and 3D and insufficient integration of 3D object localization and instruction understanding. To tackle these limitations, we first introduce NuInteract, a large-scale dataset with over 1.5M multi-view image language pairs spanning dense scene captions and diverse interactive tasks. Furthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries. The spatial processor, designed as a plug-and-play component, can be initialized with pre-trained 3D detectors to improve 3D perception. Our experiments show that DriveMonkey outperforms general LVLMs, especially achieving a 9.86% notable improvement on the 3D visual grounding task. The dataset and code will be released at https://github.com/zc-zhao/DriveMonkey.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08725
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
Zhao, Zongchuang
Fu, Haoyu
Liang, Dingkang
Zhou, Xin
Zhang, Dingyuan
Xie, Hongwei
Wang, Bing
Bai, Xiang
Computer Vision and Pattern Recognition
The Large Visual-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing research typically focuses on front-view perspectives and partial objects within scenes, struggling to achieve comprehensive scene understanding. Meanwhile, existing LVLMs suffer from the lack of mapping relationship between 2D and 3D and insufficient integration of 3D object localization and instruction understanding. To tackle these limitations, we first introduce NuInteract, a large-scale dataset with over 1.5M multi-view image language pairs spanning dense scene captions and diverse interactive tasks. Furthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries. The spatial processor, designed as a plug-and-play component, can be initialized with pre-trained 3D detectors to improve 3D perception. Our experiments show that DriveMonkey outperforms general LVLMs, especially achieving a 9.86% notable improvement on the 3D visual grounding task. The dataset and code will be released at https://github.com/zc-zhao/DriveMonkey.
title Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.08725