SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiawei, Yang, Xuan, Wang, Taiqi, Yao, Yu, Petiushko, Aleksandr, Li, Bo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908395776245760
author Zhang, Jiawei
Yang, Xuan
Wang, Taiqi
Yao, Yu
Petiushko, Aleksandr
Li, Bo
author_facet Zhang, Jiawei
Yang, Xuan
Wang, Taiqi
Yao, Yu
Petiushko, Aleksandr
Li, Bo
contents Traditional autonomous driving systems often struggle to connect high-level reasoning with low-level control, leading to suboptimal and sometimes unsafe behaviors. Recent advances in multimodal large language models (MLLMs), which process both visual and textual data, offer an opportunity to unify perception and reasoning. However, effectively embedding precise safety knowledge into MLLMs for autonomous driving remains a significant challenge. To address this, we propose SafeAuto, a framework that enhances MLLM-based autonomous driving by incorporating both unstructured and structured knowledge. First, we introduce a Position-Dependent Cross-Entropy (PDCE) loss to improve low-level control signal predictions when values are represented as text. Second, to explicitly integrate safety knowledge, we develop a reasoning component that translates traffic rules into first-order logic (e.g., "red light $\implies$ stop") and embeds them into a probabilistic graphical model (e.g., Markov Logic Network) to verify predicted actions using recognized environmental attributes. Additionally, our Multimodal Retrieval-Augmented Generation (RAG) model leverages video, control signals, and environmental attributes to learn from past driving experiences. Integrating PDCE, MLN, and Multimodal RAG, SafeAuto outperforms existing baselines across multiple datasets, enabling more accurate, reliable, and safer autonomous driving. The code is available at https://github.com/AI-secure/SafeAuto.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00211
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models
Zhang, Jiawei
Yang, Xuan
Wang, Taiqi
Yao, Yu
Petiushko, Aleksandr
Li, Bo
Robotics
Artificial Intelligence
Machine Learning
Systems and Control
Traditional autonomous driving systems often struggle to connect high-level reasoning with low-level control, leading to suboptimal and sometimes unsafe behaviors. Recent advances in multimodal large language models (MLLMs), which process both visual and textual data, offer an opportunity to unify perception and reasoning. However, effectively embedding precise safety knowledge into MLLMs for autonomous driving remains a significant challenge. To address this, we propose SafeAuto, a framework that enhances MLLM-based autonomous driving by incorporating both unstructured and structured knowledge. First, we introduce a Position-Dependent Cross-Entropy (PDCE) loss to improve low-level control signal predictions when values are represented as text. Second, to explicitly integrate safety knowledge, we develop a reasoning component that translates traffic rules into first-order logic (e.g., "red light $\implies$ stop") and embeds them into a probabilistic graphical model (e.g., Markov Logic Network) to verify predicted actions using recognized environmental attributes. Additionally, our Multimodal Retrieval-Augmented Generation (RAG) model leverages video, control signals, and environmental attributes to learn from past driving experiences. Integrating PDCE, MLN, and Multimodal RAG, SafeAuto outperforms existing baselines across multiple datasets, enabling more accurate, reliable, and safer autonomous driving. The code is available at https://github.com/AI-secure/SafeAuto.
title SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models
topic Robotics
Artificial Intelligence
Machine Learning
Systems and Control
url https://arxiv.org/abs/2503.00211