Tracking Meets Large Multimodal Models for Driving Scenario Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ishaq, Ayesha, Lahoud, Jean, Khan, Fahad Shahbaz, Khan, Salman, Cholakkal, Hisham, Anwer, Rao Muhammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913743770746880
author Ishaq, Ayesha
Lahoud, Jean
Khan, Fahad Shahbaz
Khan, Salman
Cholakkal, Hisham
Anwer, Rao Muhammad
author_facet Ishaq, Ayesha
Lahoud, Jean
Khan, Fahad Shahbaz
Khan, Salman
Cholakkal, Hisham
Anwer, Rao Muhammad
contents Large Multimodal Models (LMMs) have recently gained prominence in autonomous driving research, showcasing promising capabilities across various emerging benchmarks. LMMs specifically designed for this domain have demonstrated effective perception, planning, and prediction skills. However, many of these methods underutilize 3D spatial and temporal elements, relying mainly on image data. As a result, their effectiveness in dynamic driving environments is limited. We propose to integrate tracking information as an additional input to recover 3D spatial and temporal details that are not effectively captured in the images. We introduce a novel approach for embedding this tracking information into LMMs to enhance their spatiotemporal understanding of driving scenarios. By incorporating 3D tracking data through a track encoder, we enrich visual queries with crucial spatial and temporal cues while avoiding the computational overhead associated with processing lengthy video sequences or extensive 3D inputs. Moreover, we employ a self-supervised approach to pretrain the tracking encoder to provide LMMs with additional contextual information, significantly improving their performance in perception, planning, and prediction tasks for autonomous driving. Experimental results demonstrate the effectiveness of our approach, with a gain of 9.5% in accuracy, an increase of 7.04 points in the ChatGPT score, and 9.4% increase in the overall score over baseline models on DriveLM-nuScenes benchmark, along with a 3.7% final score improvement on DriveLM-CARLA. Our code is available at https://github.com/mbzuai-oryx/TrackingMeetsLMM
format Preprint
id arxiv_https___arxiv_org_abs_2503_14498
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tracking Meets Large Multimodal Models for Driving Scenario Understanding
Ishaq, Ayesha
Lahoud, Jean
Khan, Fahad Shahbaz
Khan, Salman
Cholakkal, Hisham
Anwer, Rao Muhammad
Computer Vision and Pattern Recognition
Robotics
Large Multimodal Models (LMMs) have recently gained prominence in autonomous driving research, showcasing promising capabilities across various emerging benchmarks. LMMs specifically designed for this domain have demonstrated effective perception, planning, and prediction skills. However, many of these methods underutilize 3D spatial and temporal elements, relying mainly on image data. As a result, their effectiveness in dynamic driving environments is limited. We propose to integrate tracking information as an additional input to recover 3D spatial and temporal details that are not effectively captured in the images. We introduce a novel approach for embedding this tracking information into LMMs to enhance their spatiotemporal understanding of driving scenarios. By incorporating 3D tracking data through a track encoder, we enrich visual queries with crucial spatial and temporal cues while avoiding the computational overhead associated with processing lengthy video sequences or extensive 3D inputs. Moreover, we employ a self-supervised approach to pretrain the tracking encoder to provide LMMs with additional contextual information, significantly improving their performance in perception, planning, and prediction tasks for autonomous driving. Experimental results demonstrate the effectiveness of our approach, with a gain of 9.5% in accuracy, an increase of 7.04 points in the ChatGPT score, and 9.4% increase in the overall score over baseline models on DriveLM-nuScenes benchmark, along with a 3.7% final score improvement on DriveLM-CARLA. Our code is available at https://github.com/mbzuai-oryx/TrackingMeetsLMM
title Tracking Meets Large Multimodal Models for Driving Scenario Understanding
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.14498