Probing Multimodal LLMs as World Models for Driving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sreeram, Shiva, Wang, Tsun-Hsuan, Maalouf, Alaa, Rosman, Guy, Karaman, Sertac, Rus, Daniela
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909365946023936
author Sreeram, Shiva
Wang, Tsun-Hsuan
Maalouf, Alaa
Rosman, Guy
Karaman, Sertac
Rus, Daniela
author_facet Sreeram, Shiva
Wang, Tsun-Hsuan
Maalouf, Alaa
Rosman, Guy
Karaman, Sertac
Rus, Daniela
contents We provide a sober look at the application of Multimodal Large Language Models (MLLMs) in autonomous driving, challenging common assumptions about their ability to interpret dynamic driving scenarios. Despite advances in models like GPT-4o, their performance in complex driving environments remains largely unexplored. Our experimental study assesses various MLLMs as world models using in-car camera perspectives and reveals that while these models excel at interpreting individual images, they struggle to synthesize coherent narratives across frames, leading to considerable inaccuracies in understanding (i) ego vehicle dynamics, (ii) interactions with other road actors, (iii) trajectory planning, and (iv) open-set scene reasoning. We introduce the Eval-LLM-Drive dataset and DriveSim simulator to enhance our evaluation, highlighting gaps in current MLLM capabilities and the need for improved models in dynamic real-world environments.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05956
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Probing Multimodal LLMs as World Models for Driving
Sreeram, Shiva
Wang, Tsun-Hsuan
Maalouf, Alaa
Rosman, Guy
Karaman, Sertac
Rus, Daniela
Robotics
Computer Vision and Pattern Recognition
We provide a sober look at the application of Multimodal Large Language Models (MLLMs) in autonomous driving, challenging common assumptions about their ability to interpret dynamic driving scenarios. Despite advances in models like GPT-4o, their performance in complex driving environments remains largely unexplored. Our experimental study assesses various MLLMs as world models using in-car camera perspectives and reveals that while these models excel at interpreting individual images, they struggle to synthesize coherent narratives across frames, leading to considerable inaccuracies in understanding (i) ego vehicle dynamics, (ii) interactions with other road actors, (iii) trajectory planning, and (iv) open-set scene reasoning. We introduce the Eval-LLM-Drive dataset and DriveSim simulator to enhance our evaluation, highlighting gaps in current MLLM capabilities and the need for improved models in dynamic real-world environments.
title Probing Multimodal LLMs as World Models for Driving
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.05956