Learning Vision-Language-Action World Models for Autonomous Driving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Guoqing, Tang, Pin, Ren, Xiangxuan, Zhao, Guodongfang, Feng, Bailan, Ma, Chao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving
von: Wang, Guoqing, et al.
Veröffentlicht: (2024)
von: Wang, Guoqing, et al.
Veröffentlicht: (2024)
SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction
von: Tang, Pin, et al.
Veröffentlicht: (2024)
von: Tang, Pin, et al.
Veröffentlicht: (2024)
VEON: Vocabulary-Enhanced Occupancy Prediction
von: Zheng, Jilai, et al.
Veröffentlicht: (2024)
von: Zheng, Jilai, et al.
Veröffentlicht: (2024)
Grounding Everything in Tokens for Multimodal Large Language Models
von: Ren, Xiangxuan, et al.
Veröffentlicht: (2025)
von: Ren, Xiangxuan, et al.
Veröffentlicht: (2025)
LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation
von: Ren, Xiangxuan, et al.
Veröffentlicht: (2025)
von: Ren, Xiangxuan, et al.
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Autonomous Driving
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
von: Jiang, Sicong, et al.
Veröffentlicht: (2025)
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026)
von: Mei, Xiaodong, et al.
Veröffentlicht: (2026)
TurboVSR: Fantastic Video Upscalers and Where to Find Them
von: Wang, Zhongdao, et al.
Veröffentlicht: (2025)
von: Wang, Zhongdao, et al.
Veröffentlicht: (2025)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
von: Yang, Zhenjie, et al.
Veröffentlicht: (2025)
von: Yang, Zhenjie, et al.
Veröffentlicht: (2025)
DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
von: Zuo, Sicheng, et al.
Veröffentlicht: (2026)
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
von: Wang, Lu, et al.
Veröffentlicht: (2025)
von: Wang, Lu, et al.
Veröffentlicht: (2025)
Vision Language Models in Autonomous Driving: A Survey and Outlook
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
HiST-VLA: A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
von: Li, Yingyan, et al.
Veröffentlicht: (2025)
von: Li, Yingyan, et al.
Veröffentlicht: (2025)
A Unified Perception-Language-Action Framework for Adaptive Autonomous Driving
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
von: Zhang, Enming, et al.
Veröffentlicht: (2025)
von: Zhang, Enming, et al.
Veröffentlicht: (2025)
VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events
von: Bhat, Mohammad Qazim, et al.
Veröffentlicht: (2026)
von: Bhat, Mohammad Qazim, et al.
Veröffentlicht: (2026)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
von: Wang, Lening, et al.
Veröffentlicht: (2024)
von: Wang, Lening, et al.
Veröffentlicht: (2024)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving
von: Xu, Zhefan, et al.
Veröffentlicht: (2026)
von: Xu, Zhefan, et al.
Veröffentlicht: (2026)
DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving
von: Fu, Yongjie, et al.
Veröffentlicht: (2024)
von: Fu, Yongjie, et al.
Veröffentlicht: (2024)
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
von: Huang, Zilin, et al.
Veröffentlicht: (2026)
von: Huang, Zilin, et al.
Veröffentlicht: (2026)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
von: Yang, Sheng, et al.
Veröffentlicht: (2025)
von: Yang, Sheng, et al.
Veröffentlicht: (2025)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models
von: Wang, Yujin, et al.
Veröffentlicht: (2024)
von: Wang, Yujin, et al.
Veröffentlicht: (2024)
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
von: Sha, Lin, et al.
Veröffentlicht: (2026)
von: Sha, Lin, et al.
Veröffentlicht: (2026)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
von: Wang, Zhaohui, et al.
Veröffentlicht: (2025)
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
von: He, Jingtao, et al.
Veröffentlicht: (2026)
von: He, Jingtao, et al.
Veröffentlicht: (2026)
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
von: Huang, Minqing, et al.
Veröffentlicht: (2026)
von: Huang, Minqing, et al.
Veröffentlicht: (2026)
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving
von: Zheng, Peiru, et al.
Veröffentlicht: (2024)
von: Zheng, Peiru, et al.
Veröffentlicht: (2024)
VECTOR-Drive: Tightly Coupled Vision-Language and Trajectory Expert Routing for End-to-End Autonomous Driving
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
von: Zhao, Rui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving
von: Wang, Guoqing, et al.
Veröffentlicht: (2024) -
SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy Prediction
von: Tang, Pin, et al.
Veröffentlicht: (2024) -
VEON: Vocabulary-Enhanced Occupancy Prediction
von: Zheng, Jilai, et al.
Veröffentlicht: (2024) -
Grounding Everything in Tokens for Multimodal Large Language Models
von: Ren, Xiangxuan, et al.
Veröffentlicht: (2025) -
LiteFusion: Taming 3D Object Detectors from Vision-Based to Multi-Modal with Minimal Adaptation
von: Ren, Xiangxuan, et al.
Veröffentlicht: (2025)