dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Yingzi, Cao, Yulong, Ding, Wenhao, Zhang, Shuibai, Wang, Yan, Ivanovic, Boris, Jiang, Ming, Pavone, Marco, Xiao, Chaowei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RealDrive: Retrieval-Augmented Driving with Diffusion Models
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
RealGen: Retrieval Augmented Generation for Controllable Traffic Scenarios
von: Ding, Wenhao, et al.
Veröffentlicht: (2023)
von: Ding, Wenhao, et al.
Veröffentlicht: (2023)
Gen-Drive: Enhancing Diffusion Generative Driving Policies with Reward Modeling and Reinforcement Learning Fine-tuning
von: Huang, Zhiyu, et al.
Veröffentlicht: (2024)
von: Huang, Zhiyu, et al.
Veröffentlicht: (2024)
Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
von: Wu, Chengyue, et al.
Veröffentlicht: (2026)
von: Wu, Chengyue, et al.
Veröffentlicht: (2026)
Surprise Potential as a Measure of Interactivity in Driving Scenarios
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
von: Li, Nanxi, et al.
Veröffentlicht: (2025)
von: Li, Nanxi, et al.
Veröffentlicht: (2025)
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
von: Zhao, Zhengyue, et al.
Veröffentlicht: (2025)
von: Zhao, Zhengyue, et al.
Veröffentlicht: (2025)
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
DreamDrive: Generative 4D Scene Modeling from Street View Images
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
von: Mao, Jiageng, et al.
Veröffentlicht: (2024)
Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
von: Ivanovic, Boris, et al.
Veröffentlicht: (2025)
von: Ivanovic, Boris, et al.
Veröffentlicht: (2025)
Latent Chain-of-Thought World Modeling for End-to-End Driving
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models
von: Zhang, Zhejun, et al.
Veröffentlicht: (2024)
von: Zhang, Zhejun, et al.
Veröffentlicht: (2024)
DTPP: Differentiable Joint Conditional Prediction and Cost Evaluation for Tree Policy Planning in Autonomous Driving
von: Huang, Zhiyu, et al.
Veröffentlicht: (2023)
von: Huang, Zhiyu, et al.
Veröffentlicht: (2023)
Driving Everywhere with Large Language Model Policy Adaptation
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
Promptable Closed-loop Traffic Simulation
von: Tan, Shuhan, et al.
Veröffentlicht: (2024)
von: Tan, Shuhan, et al.
Veröffentlicht: (2024)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
von: Gu, Xunjiang, et al.
Veröffentlicht: (2024)
von: Gu, Xunjiang, et al.
Veröffentlicht: (2024)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
von: Gu, Xunjiang, et al.
Veröffentlicht: (2024)
von: Gu, Xunjiang, et al.
Veröffentlicht: (2024)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving
von: Tian, Ran, et al.
Veröffentlicht: (2024)
von: Tian, Ran, et al.
Veröffentlicht: (2024)
Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators
von: Chakraborty, Kaustav, et al.
Veröffentlicht: (2025)
von: Chakraborty, Kaustav, et al.
Veröffentlicht: (2025)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
von: Yang, Jiawei, et al.
Veröffentlicht: (2025)
von: Yang, Jiawei, et al.
Veröffentlicht: (2025)
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025)
von: Garcia-Cobo, Guillermo, et al.
Veröffentlicht: (2025)
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
von: Ma, Yingzi, et al.
Veröffentlicht: (2026)
von: Ma, Yingzi, et al.
Veröffentlicht: (2026)
Data Scaling Laws for End-to-End Autonomous Driving
von: Naumann, Alexander, et al.
Veröffentlicht: (2025)
von: Naumann, Alexander, et al.
Veröffentlicht: (2025)
Accelerating Structured Chain-of-Thought in Autonomous Vehicles
von: Gu, Yi, et al.
Veröffentlicht: (2026)
von: Gu, Yi, et al.
Veröffentlicht: (2026)
Language-Image Models with 3D Understanding
von: Cho, Jang Hyun, et al.
Veröffentlicht: (2024)
von: Cho, Jang Hyun, et al.
Veröffentlicht: (2024)
Pseudo-Simulation for Autonomous Driving
von: Cao, Wei, et al.
Veröffentlicht: (2025)
von: Cao, Wei, et al.
Veröffentlicht: (2025)
AppleVLM: End-to-end Autonomous Driving with Advanced Perception and Planning-Enhanced Vision-Language Models
von: Han, Yuxuan, et al.
Veröffentlicht: (2026)
von: Han, Yuxuan, et al.
Veröffentlicht: (2026)
WIPI: A New Web Threat for LLM-Driven Web Agents
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2025)
System-Level Safety Monitoring and Recovery for Perception Failures in Autonomous Vehicles
von: Chakraborty, Kaustav, et al.
Veröffentlicht: (2024)
von: Chakraborty, Kaustav, et al.
Veröffentlicht: (2024)
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
von: Lu, Ziqi, et al.
Veröffentlicht: (2024)
von: Lu, Ziqi, et al.
Veröffentlicht: (2024)
Parallelized Spatiotemporal Binding
von: Singh, Gautam, et al.
Veröffentlicht: (2024)
von: Singh, Gautam, et al.
Veröffentlicht: (2024)
Sample-Efficient Safety Assurances using Conformal Prediction
von: Luo, Rachel, et al.
Veröffentlicht: (2021)
von: Luo, Rachel, et al.
Veröffentlicht: (2021)
VLM-UDMC: VLM-Enhanced Unified Decision-Making and Motion Control for Urban Autonomous Driving
von: Liu, Haichao, et al.
Veröffentlicht: (2025)
von: Liu, Haichao, et al.
Veröffentlicht: (2025)
VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving
von: Long, Keke, et al.
Veröffentlicht: (2024)
von: Long, Keke, et al.
Veröffentlicht: (2024)
ExpertAD: Enhancing Autonomous Driving Systems with Mixture of Experts
von: Jiang, Haowen, et al.
Veröffentlicht: (2025)
von: Jiang, Haowen, et al.
Veröffentlicht: (2025)
FloorplanVLM: A Vision-Language Model for Floorplan Vectorization
von: Liu, Yuanqing, et al.
Veröffentlicht: (2026)
von: Liu, Yuanqing, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RealDrive: Retrieval-Augmented Driving with Diffusion Models
von: Ding, Wenhao, et al.
Veröffentlicht: (2025) -
RealGen: Retrieval Augmented Generation for Controllable Traffic Scenarios
von: Ding, Wenhao, et al.
Veröffentlicht: (2023) -
Gen-Drive: Enhancing Diffusion Generative Driving Policies with Reward Modeling and Reinforcement Learning Fine-tuning
von: Huang, Zhiyu, et al.
Veröffentlicht: (2024) -
Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
von: Wu, Chengyue, et al.
Veröffentlicht: (2026) -
Surprise Potential as a Measure of Interactivity in Driving Scenarios
von: Ding, Wenhao, et al.
Veröffentlicht: (2025)