AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xing, Shuo, Hua, Hongyuan, Gao, Xiangbo, Zhu, Shenzhe, Li, Renjie, Tian, Kexin, Li, Xiaopeng, Huang, Heng, Yang, Tianbao, Wang, Zhangyang, Zhou, Yang, Yao, Huaxiu, Tu, Zhengzhong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917179175206912
author Xing, Shuo
Hua, Hongyuan
Gao, Xiangbo
Zhu, Shenzhe
Li, Renjie
Tian, Kexin
Li, Xiaopeng
Huang, Heng
Yang, Tianbao
Wang, Zhangyang
Zhou, Yang
Yao, Huaxiu
Tu, Zhengzhong
author_facet Xing, Shuo
Hua, Hongyuan
Gao, Xiangbo
Zhu, Shenzhe
Li, Renjie
Tian, Kexin
Li, Xiaopeng
Huang, Heng
Yang, Tianbao
Wang, Zhangyang
Zhou, Yang
Yao, Huaxiu
Tu, Zhengzhong
contents Recent advancements in large vision language models (VLMs) tailored for autonomous driving (AD) have shown strong scene understanding and reasoning capabilities, making them undeniable candidates for end-to-end driving systems. However, limited work exists on studying the trustworthiness of DriveVLMs -- a critical factor that directly impacts public transportation safety. In this paper, we introduce AutoTrust, a comprehensive trustworthiness benchmark for large vision-language models in autonomous driving (DriveVLMs), considering diverse perspectives -- including trustfulness, safety, robustness, privacy, and fairness. We constructed the largest visual question-answering dataset for investigating trustworthiness issues in driving scenarios, comprising over 10k unique scenes and 18k queries. We evaluated six publicly available VLMs, spanning from generalist to specialist, from open-source to commercial models. Our exhaustive evaluations have unveiled previously undiscovered vulnerabilities of DriveVLMs to trustworthiness threats. Specifically, we found that the general VLMs like LLaVA-v1.6 and GPT-4o-mini surprisingly outperform specialized models fine-tuned for driving in terms of overall trustworthiness. DriveVLMs like DriveLM-Agent are particularly vulnerable to disclosing sensitive information. Additionally, both generalist and specialist VLMs remain susceptible to adversarial attacks and struggle to ensure unbiased decision-making across diverse environments and populations. Our findings call for immediate and decisive action to address the trustworthiness of DriveVLMs -- an issue of critical importance to public safety and the welfare of all citizens relying on autonomous transportation systems. We release all the codes and datasets in https://github.com/taco-group/AutoTrust.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15206
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
Xing, Shuo
Hua, Hongyuan
Gao, Xiangbo
Zhu, Shenzhe
Li, Renjie
Tian, Kexin
Li, Xiaopeng
Huang, Heng
Yang, Tianbao
Wang, Zhangyang
Zhou, Yang
Yao, Huaxiu
Tu, Zhengzhong
Computer Vision and Pattern Recognition
Machine Learning
Robotics
Recent advancements in large vision language models (VLMs) tailored for autonomous driving (AD) have shown strong scene understanding and reasoning capabilities, making them undeniable candidates for end-to-end driving systems. However, limited work exists on studying the trustworthiness of DriveVLMs -- a critical factor that directly impacts public transportation safety. In this paper, we introduce AutoTrust, a comprehensive trustworthiness benchmark for large vision-language models in autonomous driving (DriveVLMs), considering diverse perspectives -- including trustfulness, safety, robustness, privacy, and fairness. We constructed the largest visual question-answering dataset for investigating trustworthiness issues in driving scenarios, comprising over 10k unique scenes and 18k queries. We evaluated six publicly available VLMs, spanning from generalist to specialist, from open-source to commercial models. Our exhaustive evaluations have unveiled previously undiscovered vulnerabilities of DriveVLMs to trustworthiness threats. Specifically, we found that the general VLMs like LLaVA-v1.6 and GPT-4o-mini surprisingly outperform specialized models fine-tuned for driving in terms of overall trustworthiness. DriveVLMs like DriveLM-Agent are particularly vulnerable to disclosing sensitive information. Additionally, both generalist and specialist VLMs remain susceptible to adversarial attacks and struggle to ensure unbiased decision-making across diverse environments and populations. Our findings call for immediate and decisive action to address the trustworthiness of DriveVLMs -- an issue of critical importance to public safety and the welfare of all citizens relying on autonomous transportation systems. We release all the codes and datasets in https://github.com/taco-group/AutoTrust.
title AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
topic Computer Vision and Pattern Recognition
Machine Learning
Robotics
url https://arxiv.org/abs/2412.15206