DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Erfei, Wang, Wenhai, Li, Zhiqi, Xie, Jiangwei, Zou, Haoming, Deng, Hanming, Luo, Gen, Lu, Lewei, Zhu, Xizhou, Dai, Jifeng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
by: Wu, Zehuan, et al.
Published: (2024)
by: Wu, Zehuan, et al.
Published: (2024)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
by: Tian, Changyao, et al.
Published: (2025)
by: Tian, Changyao, et al.
Published: (2025)
M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving
by: Xu, Dongyang, et al.
Published: (2024)
by: Xu, Dongyang, et al.
Published: (2024)
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
by: Wu, Yanhao, et al.
Published: (2026)
by: Wu, Yanhao, et al.
Published: (2026)
CoMemo: LVLMs Need Image Context with Image Memory
by: Liu, Shi, et al.
Published: (2025)
by: Liu, Shi, et al.
Published: (2025)
Characterized Diffusion Networks for Enhanced Autonomous Driving Trajectory Prediction
by: Li, Haoming
Published: (2024)
by: Li, Haoming
Published: (2024)
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
by: Luo, Gen, et al.
Published: (2025)
by: Luo, Gen, et al.
Published: (2025)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
by: Liu, Yangzhou, et al.
Published: (2024)
by: Liu, Yangzhou, et al.
Published: (2024)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
by: Gao, Zhangwei, et al.
Published: (2024)
by: Gao, Zhangwei, et al.
Published: (2024)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
by: Xiong, Yuwen, et al.
Published: (2024)
by: Xiong, Yuwen, et al.
Published: (2024)
Belief State Planning for Autonomous Driving
by: Hubmann, Constantin
Published: (2021)
by: Hubmann, Constantin
Published: (2021)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
by: Chen, Zhe, et al.
Published: (2023)
by: Chen, Zhe, et al.
Published: (2023)
Driving Intents Amplify Planning-Oriented Reinforcement Learning
by: Lu, Hengtong, et al.
Published: (2026)
by: Lu, Hengtong, et al.
Published: (2026)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
by: Duan, Yuchen, et al.
Published: (2024)
by: Duan, Yuchen, et al.
Published: (2024)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
by: Yang, Chenyu, et al.
Published: (2024)
by: Yang, Chenyu, et al.
Published: (2024)
HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving
by: Chen, Zhiwen, et al.
Published: (2025)
by: Chen, Zhiwen, et al.
Published: (2025)
ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
by: Li, Jingyu, et al.
Published: (2025)
by: Li, Jingyu, et al.
Published: (2025)
Driving with InternVL: Oustanding Champion in the Track on Driving with Language of the Autonomous Grand Challenge at CVPR 2024
by: Li, Jiahan, et al.
Published: (2024)
by: Li, Jiahan, et al.
Published: (2024)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
by: Wang, Zhaokai, et al.
Published: (2025)
by: Wang, Zhaokai, et al.
Published: (2025)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
by: Wang, Weiyun, et al.
Published: (2024)
by: Wang, Weiyun, et al.
Published: (2024)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
by: Tao, Chenxin, et al.
Published: (2024)
by: Tao, Chenxin, et al.
Published: (2024)
Driving-Video Dehazing with Non-Aligned Regularization for Safety Assistance
by: Fan, Junkai, et al.
Published: (2024)
by: Fan, Junkai, et al.
Published: (2024)
Pullback Attractors for Nonautonomous Reaction–Diffusion Equations With the Driving Delay Term in ℝN
by: Yong Ren, et al.
Published: (2025)
by: Yong Ren, et al.
Published: (2025)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving
by: Xiong, Zhexiao, et al.
Published: (2026)
by: Xiong, Zhexiao, et al.
Published: (2026)
LAP: Fast LAtent Diffusion Planner for Autonomous Driving
by: Zhang, Jinhao, et al.
Published: (2025)
by: Zhang, Jinhao, et al.
Published: (2025)
RealDriveSim: A Realistic Multi-Modal Multi-Task Synthetic Dataset for Autonomous Driving
by: Jadon, Arpit, et al.
Published: (2025)
by: Jadon, Arpit, et al.
Published: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
by: Wu, Jiannan, et al.
Published: (2024)
by: Wu, Jiannan, et al.
Published: (2024)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
by: Luo, Gen, et al.
Published: (2024)
by: Luo, Gen, et al.
Published: (2024)
EqDrive: Efficient Equivariant Motion Forecasting with Multi-Modality for Autonomous Driving
by: Wang, Yuping, et al.
Published: (2023)
by: Wang, Yuping, et al.
Published: (2023)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
by: Yang, Chenyu, et al.
Published: (2024)
by: Yang, Chenyu, et al.
Published: (2024)
Demystify Transformers & Convolutions in Modern Image Deep Networks
by: Hu, Xiaowei, et al.
Published: (2022)
by: Hu, Xiaowei, et al.
Published: (2022)
Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior Modeling
by: Tan, Tianyi, et al.
Published: (2025)
by: Tan, Tianyi, et al.
Published: (2025)
Cognitive-Hierarchy Guided End-to-End Planning for Autonomous Driving
by: Wang, Zhennan, et al.
Published: (2025)
by: Wang, Zhennan, et al.
Published: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
by: Luo, Gen, et al.
Published: (2025)
by: Luo, Gen, et al.
Published: (2025)
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving
by: Huang, Yuzhou, et al.
Published: (2026)
by: Huang, Yuzhou, et al.
Published: (2026)
Improving Agent Behaviors with RL Fine-tuning for Autonomous Driving
by: Peng, Zhenghao, et al.
Published: (2024)
by: Peng, Zhenghao, et al.
Published: (2024)
Align2Act: Instruction-Tuned Models for Human-Aligned Autonomous Driving
by: Jaisankar, Kanishkha, et al.
Published: (2025)
by: Jaisankar, Kanishkha, et al.
Published: (2025)
SEIDM: A Safe and Efficient Intelligent Driver Model for Autonomous Driving Behavior
by: Yao, Yuyang, et al.
Published: (2026)
by: Yao, Yuyang, et al.
Published: (2026)
Similar Items
-
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
by: Wu, Zehuan, et al.
Published: (2024) -
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
by: Tian, Changyao, et al.
Published: (2025) -
M2DA: Multi-Modal Fusion Transformer Incorporating Driver Attention for Autonomous Driving
by: Xu, Dongyang, et al.
Published: (2024) -
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
by: Wu, Yanhao, et al.
Published: (2026) -
CoMemo: LVLMs Need Image Context with Image Memory
by: Liu, Shi, et al.
Published: (2025)