DriveMind: A Dual Visual Language Model-based Reinforcement Learning Framework for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wasif, Dawood, Moore, Terrence J., Reddy, Chandan K., Free-Nelson, Frederica, Yoon, Seunghyun, Lim, Hyuk, Kim, Dan Dongseong, Cho, Jin-Hee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912963907026944
author Wasif, Dawood
Moore, Terrence J.
Reddy, Chandan K.
Free-Nelson, Frederica
Yoon, Seunghyun
Lim, Hyuk
Kim, Dan Dongseong
Cho, Jin-Hee
author_facet Wasif, Dawood
Moore, Terrence J.
Reddy, Chandan K.
Free-Nelson, Frederica
Yoon, Seunghyun
Lim, Hyuk
Kim, Dan Dongseong
Cho, Jin-Hee
contents End-to-end autonomous driving systems map sensor data directly to control commands, but remain opaque, lack interpretability, and offer no formal safety guarantees. While recent vision-language-guided reinforcement learning (RL) methods introduce semantic feedback, they often rely on static prompts and fixed objectives, limiting adaptability to dynamic driving scenes. We present DriveMind, a unified semantic reward framework that integrates: (i) a contrastive Vision-Language Model (VLM) encoder for stepwise semantic anchoring; (ii) a novelty-triggered VLM encoder-decoder, fine-tuned via chain-of-thought (CoT) distillation, for dynamic prompt generation upon semantic drift; (iii) a hierarchical safety module enforcing kinematic constraints (e.g., speed, lane centering, stability); and (iv) a compact predictive world model to reward alignment with anticipated ideal states. DriveMind achieves 19.4 +/- 2.3 km/h average speed, 0.98 +/- 0.03 route completion, and near-zero collisions in CARLA Town 2, outperforming baselines by over 4% in success rate. Its semantic reward generalizes zero-shot to real dash-cam data with minimal distributional shift, demonstrating robust cross-domain alignment and potential for real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00819
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DriveMind: A Dual Visual Language Model-based Reinforcement Learning Framework for Autonomous Driving
Wasif, Dawood
Moore, Terrence J.
Reddy, Chandan K.
Free-Nelson, Frederica
Yoon, Seunghyun
Lim, Hyuk
Kim, Dan Dongseong
Cho, Jin-Hee
Robotics
Artificial Intelligence
End-to-end autonomous driving systems map sensor data directly to control commands, but remain opaque, lack interpretability, and offer no formal safety guarantees. While recent vision-language-guided reinforcement learning (RL) methods introduce semantic feedback, they often rely on static prompts and fixed objectives, limiting adaptability to dynamic driving scenes. We present DriveMind, a unified semantic reward framework that integrates: (i) a contrastive Vision-Language Model (VLM) encoder for stepwise semantic anchoring; (ii) a novelty-triggered VLM encoder-decoder, fine-tuned via chain-of-thought (CoT) distillation, for dynamic prompt generation upon semantic drift; (iii) a hierarchical safety module enforcing kinematic constraints (e.g., speed, lane centering, stability); and (iv) a compact predictive world model to reward alignment with anticipated ideal states. DriveMind achieves 19.4 +/- 2.3 km/h average speed, 0.98 +/- 0.03 route completion, and near-zero collisions in CARLA Town 2, outperforming baselines by over 4% in success rate. Its semantic reward generalizes zero-shot to real dash-cam data with minimal distributional shift, demonstrating robust cross-domain alignment and potential for real-world deployment.
title DriveMind: A Dual Visual Language Model-based Reinforcement Learning Framework for Autonomous Driving
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2506.00819