DriveMind: A Dual Visual Language Model-based Reinforcement Learning Framework for Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912963907026944 |
|---|---|
| author | Wasif, Dawood Moore, Terrence J. Reddy, Chandan K. Free-Nelson, Frederica Yoon, Seunghyun Lim, Hyuk Kim, Dan Dongseong Cho, Jin-Hee |
| author_facet | Wasif, Dawood Moore, Terrence J. Reddy, Chandan K. Free-Nelson, Frederica Yoon, Seunghyun Lim, Hyuk Kim, Dan Dongseong Cho, Jin-Hee |
| contents | End-to-end autonomous driving systems map sensor data directly to control commands, but remain opaque, lack interpretability, and offer no formal safety guarantees. While recent vision-language-guided reinforcement learning (RL) methods introduce semantic feedback, they often rely on static prompts and fixed objectives, limiting adaptability to dynamic driving scenes. We present DriveMind, a unified semantic reward framework that integrates: (i) a contrastive Vision-Language Model (VLM) encoder for stepwise semantic anchoring; (ii) a novelty-triggered VLM encoder-decoder, fine-tuned via chain-of-thought (CoT) distillation, for dynamic prompt generation upon semantic drift; (iii) a hierarchical safety module enforcing kinematic constraints (e.g., speed, lane centering, stability); and (iv) a compact predictive world model to reward alignment with anticipated ideal states. DriveMind achieves 19.4 +/- 2.3 km/h average speed, 0.98 +/- 0.03 route completion, and near-zero collisions in CARLA Town 2, outperforming baselines by over 4% in success rate. Its semantic reward generalizes zero-shot to real dash-cam data with minimal distributional shift, demonstrating robust cross-domain alignment and potential for real-world deployment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_00819 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DriveMind: A Dual Visual Language Model-based Reinforcement Learning Framework for Autonomous Driving Wasif, Dawood Moore, Terrence J. Reddy, Chandan K. Free-Nelson, Frederica Yoon, Seunghyun Lim, Hyuk Kim, Dan Dongseong Cho, Jin-Hee Robotics Artificial Intelligence End-to-end autonomous driving systems map sensor data directly to control commands, but remain opaque, lack interpretability, and offer no formal safety guarantees. While recent vision-language-guided reinforcement learning (RL) methods introduce semantic feedback, they often rely on static prompts and fixed objectives, limiting adaptability to dynamic driving scenes. We present DriveMind, a unified semantic reward framework that integrates: (i) a contrastive Vision-Language Model (VLM) encoder for stepwise semantic anchoring; (ii) a novelty-triggered VLM encoder-decoder, fine-tuned via chain-of-thought (CoT) distillation, for dynamic prompt generation upon semantic drift; (iii) a hierarchical safety module enforcing kinematic constraints (e.g., speed, lane centering, stability); and (iv) a compact predictive world model to reward alignment with anticipated ideal states. DriveMind achieves 19.4 +/- 2.3 km/h average speed, 0.98 +/- 0.03 route completion, and near-zero collisions in CARLA Town 2, outperforming baselines by over 4% in success rate. Its semantic reward generalizes zero-shot to real dash-cam data with minimal distributional shift, demonstrating robust cross-domain alignment and potential for real-world deployment. |
| title | DriveMind: A Dual Visual Language Model-based Reinforcement Learning Framework for Autonomous Driving |
| topic | Robotics Artificial Intelligence |
| url | https://arxiv.org/abs/2506.00819 |