Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Linfeng, Sun, Yiming, Wu, Sihao, Liu, Jiaxu, Huang, Xiaowei
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913575235223552
author He, Linfeng
Sun, Yiming
Wu, Sihao
Liu, Jiaxu
Huang, Xiaowei
author_facet He, Linfeng
Sun, Yiming
Wu, Sihao
Liu, Jiaxu
Huang, Xiaowei
contents In this paper, we propose a novel framework for enhancing visual comprehension in autonomous driving systems by integrating visual language models (VLMs) with additional visual perception module specialised in object detection. We extend the Llama-Adapter architecture by incorporating a YOLOS-based detection network alongside the CLIP perception network, addressing limitations in object detection and localisation. Our approach introduces camera ID-separators to improve multi-view processing, crucial for comprehensive environmental awareness. Experiments on the DriveLM visual question answering challenge demonstrate significant improvements over baseline models, with enhanced performance in ChatGPT scores, BLEU scores, and CIDEr metrics, indicating closeness of model answer to ground truth. Our method represents a promising step towards more capable and interpretable autonomous driving systems. Possible safety enhancement enabled by detection modality is also discussed.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05898
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
He, Linfeng
Sun, Yiming
Wu, Sihao
Liu, Jiaxu
Huang, Xiaowei
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
In this paper, we propose a novel framework for enhancing visual comprehension in autonomous driving systems by integrating visual language models (VLMs) with additional visual perception module specialised in object detection. We extend the Llama-Adapter architecture by incorporating a YOLOS-based detection network alongside the CLIP perception network, addressing limitations in object detection and localisation. Our approach introduces camera ID-separators to improve multi-view processing, crucial for comprehensive environmental awareness. Experiments on the DriveLM visual question answering challenge demonstrate significant improvements over baseline models, with enhanced performance in ChatGPT scores, BLEU scores, and CIDEr metrics, indicating closeness of model answer to ground truth. Our method represents a promising step towards more capable and interpretable autonomous driving systems. Possible safety enhancement enabled by detection modality is also discussed.
title Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2411.05898