RoboDriveVLM: A Novel Benchmark and Baseline towards Robust Vision-Language Models for Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Dacheng, Qi, Mengshi, Shu, Peng, Zhang, Zhining, Lin, Yuxin, Liu, Liang, Ma, Huadong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
by: Lin, Yuxin, et al.
Published: (2025)
by: Lin, Yuxin, et al.
Published: (2025)
Improving Batch Normalization with TTA for Robust Object Detection in Self-Driving
by: Liao, Dacheng, et al.
Published: (2024)
by: Liao, Dacheng, et al.
Published: (2024)
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving
by: Lv, Changsheng, et al.
Published: (2024)
by: Lv, Changsheng, et al.
Published: (2024)
SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation
by: Ye, Hao, et al.
Published: (2025)
by: Ye, Hao, et al.
Published: (2025)
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
by: Tian, Xiaoyu, et al.
Published: (2024)
by: Tian, Xiaoyu, et al.
Published: (2024)
Towards Efficient Object Re-Identification with A Novel Cloud-Edge Collaborative Framework
by: Wang, Chuanming, et al.
Published: (2024)
by: Wang, Chuanming, et al.
Published: (2024)
Learning Group Interactions and Semantic Intentions for Multi-Object Trajectory Prediction
by: Qi, Mengshi, et al.
Published: (2024)
by: Qi, Mengshi, et al.
Published: (2024)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Uncovering the human motion pattern: Pattern Memory-based Diffusion Model for Trajectory Prediction
by: Yang, Yuxin, et al.
Published: (2024)
by: Yang, Yuxin, et al.
Published: (2024)
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events
by: Bhat, Mohammad Qazim, et al.
Published: (2026)
by: Bhat, Mohammad Qazim, et al.
Published: (2026)
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
Multi-Stage Contrastive Regression for Action Quality Assessment
by: An, Qi, et al.
Published: (2024)
by: An, Qi, et al.
Published: (2024)
DriVLM: Domain Adaptation of Vision-Language Models in Autonomous Driving
by: Zheng, Xuran, et al.
Published: (2025)
by: Zheng, Xuran, et al.
Published: (2025)
DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving
by: Fu, Yongjie, et al.
Published: (2024)
by: Fu, Yongjie, et al.
Published: (2024)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
by: Wang, Lu, et al.
Published: (2025)
by: Wang, Lu, et al.
Published: (2025)
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
by: Huang, Zilin, et al.
Published: (2026)
by: Huang, Zilin, et al.
Published: (2026)
VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving
by: Long, Keke, et al.
Published: (2024)
by: Long, Keke, et al.
Published: (2024)
LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
by: Jiao, Siwen, et al.
Published: (2024)
by: Jiao, Siwen, et al.
Published: (2024)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
by: Zhong, Yaoyao, et al.
Published: (2023)
by: Zhong, Yaoyao, et al.
Published: (2023)
A Hierarchical Test Platform for Vision Language Model (VLM)-Integrated Real-World Autonomous Driving
by: Zhou, Yupeng, et al.
Published: (2025)
by: Zhou, Yupeng, et al.
Published: (2025)
AppleVLM: End-to-end Autonomous Driving with Advanced Perception and Planning-Enhanced Vision-Language Models
by: Han, Yuxuan, et al.
Published: (2026)
by: Han, Yuxuan, et al.
Published: (2026)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
Human Grasp Generation for Rigid and Deformable Objects with Decomposed VQ-VAE
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026)
by: Ou, Yilin, et al.
Published: (2026)
Decomposed Vector-Quantized Variational Autoencoder for Human Grasp Generation
by: Zhao, Zhe, et al.
Published: (2024)
by: Zhao, Zhe, et al.
Published: (2024)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving
by: Li, Jiaxiang, et al.
Published: (2026)
by: Li, Jiaxiang, et al.
Published: (2026)
Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Semi-Supervised Teacher-Reference-Student Architecture for Action Quality Assessment
by: Yun, Wulian, et al.
Published: (2024)
by: Yun, Wulian, et al.
Published: (2024)
A New Teacher-Reviewer-Student Framework for Semi-supervised 2D Human Pose Estimation
by: Yun, Wulian, et al.
Published: (2025)
by: Yun, Wulian, et al.
Published: (2025)
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models
by: Zhou, Xirui, et al.
Published: (2025)
by: Zhou, Xirui, et al.
Published: (2025)
Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models
by: Jia, Xiaosong, et al.
Published: (2026)
by: Jia, Xiaosong, et al.
Published: (2026)
Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
by: Liao, Yifan, et al.
Published: (2025)
by: Liao, Yifan, et al.
Published: (2025)
Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph Generation
by: Lv, Changsheng, et al.
Published: (2025)
by: Lv, Changsheng, et al.
Published: (2025)
The Constant Eye: Benchmarking and Bridging Appearance Robustness in Autonomous Driving
by: Wang, Jiabao, et al.
Published: (2026)
by: Wang, Jiabao, et al.
Published: (2026)
Similar Items
-
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
by: Lin, Yuxin, et al.
Published: (2025) -
Improving Batch Normalization with TTA for Robust Object Detection in Self-Driving
by: Liao, Dacheng, et al.
Published: (2024) -
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
by: Qi, Mengshi, et al.
Published: (2025) -
T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving
by: Lv, Changsheng, et al.
Published: (2024) -
SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation
by: Ye, Hao, et al.
Published: (2025)