Is Your VLM for Autonomous Driving Safety-Ready? A Comprehensive Benchmark for Evaluating External and In-Cabin Risks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Meng, Xianhui, Zhang, Yuchen, Huang, Zhijian, Lu, Zheng, Ji, Ziling, Yin, Yaoyao, Zhang, Hongyuan, Jiang, Guangfeng, Lin, Yandan, Chen, Long, Ye, Hangjun, Zhang, Li, Liu, Jun, Hao, Xiaoshuai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909913283821568
author Meng, Xianhui
Zhang, Yuchen
Huang, Zhijian
Lu, Zheng
Ji, Ziling
Yin, Yaoyao
Zhang, Hongyuan
Jiang, Guangfeng
Lin, Yandan
Chen, Long
Ye, Hangjun
Zhang, Li
Liu, Jun
Hao, Xiaoshuai
author_facet Meng, Xianhui
Zhang, Yuchen
Huang, Zhijian
Lu, Zheng
Ji, Ziling
Yin, Yaoyao
Zhang, Hongyuan
Jiang, Guangfeng
Lin, Yandan
Chen, Long
Ye, Hangjun
Zhang, Li
Liu, Jun
Hao, Xiaoshuai
contents Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns. This issue arises from the lack of comprehensive benchmarks that assess both external environmental risks and in-cabin driving behavior safety simultaneously. To bridge this critical gap, we introduce DSBench, the first comprehensive Driving Safety Benchmark designed to assess a VLM's awareness of various safety risks in a unified manner. DSBench encompasses two major categories: external environmental risks and in-cabin driving behavior safety, divided into 10 key categories and a total of 28 sub-categories. This comprehensive evaluation covers a wide range of scenarios, ensuring a thorough assessment of VLMs' performance in safety-critical contexts. Extensive evaluations across various mainstream open-source and closed-source VLMs reveal significant performance degradation under complex safety-critical situations, highlighting urgent safety concerns. To address this, we constructed a large dataset of 98K instances focused on in-cabin and external safety scenarios, showing that fine-tuning on this dataset significantly enhances the safety performance of existing VLMs and paves the way for advancing autonomous driving technology. The benchmark toolkit, code, and model checkpoints will be publicly accessible.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14592
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is Your VLM for Autonomous Driving Safety-Ready? A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
Meng, Xianhui
Zhang, Yuchen
Huang, Zhijian
Lu, Zheng
Ji, Ziling
Yin, Yaoyao
Zhang, Hongyuan
Jiang, Guangfeng
Lin, Yandan
Chen, Long
Ye, Hangjun
Zhang, Li
Liu, Jun
Hao, Xiaoshuai
Robotics
Artificial Intelligence
Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns. This issue arises from the lack of comprehensive benchmarks that assess both external environmental risks and in-cabin driving behavior safety simultaneously. To bridge this critical gap, we introduce DSBench, the first comprehensive Driving Safety Benchmark designed to assess a VLM's awareness of various safety risks in a unified manner. DSBench encompasses two major categories: external environmental risks and in-cabin driving behavior safety, divided into 10 key categories and a total of 28 sub-categories. This comprehensive evaluation covers a wide range of scenarios, ensuring a thorough assessment of VLMs' performance in safety-critical contexts. Extensive evaluations across various mainstream open-source and closed-source VLMs reveal significant performance degradation under complex safety-critical situations, highlighting urgent safety concerns. To address this, we constructed a large dataset of 98K instances focused on in-cabin and external safety scenarios, showing that fine-tuning on this dataset significantly enhances the safety performance of existing VLMs and paves the way for advancing autonomous driving technology. The benchmark toolkit, code, and model checkpoints will be publicly accessible.
title Is Your VLM for Autonomous Driving Safety-Ready? A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2511.14592