OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yiwei, Chen, Xuesong, Gao, Jin, Wang, Hanshi, Ge, Fudong, Hu, Weiming, Shi, Shaoshuai, Zhang, Zhipeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Integrating Diverse Assignment Strategies into DETRs
by: Zhang, Yiwei, et al.
Published: (2026)
by: Zhang, Yiwei, et al.
Published: (2026)
MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection
by: Wang, Hanshi, et al.
Published: (2025)
by: Wang, Hanshi, et al.
Published: (2025)
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
by: Chen, Xuesong, et al.
Published: (2025)
by: Chen, Xuesong, et al.
Published: (2025)
Online Segment Any 3D Thing as Instance Tracking
by: Wang, Hanshi, et al.
Published: (2025)
by: Wang, Hanshi, et al.
Published: (2025)
DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving
by: Shi, Chen, et al.
Published: (2026)
by: Shi, Chen, et al.
Published: (2026)
BEV$^2$PR: BEV-Enhanced Visual Place Recognition with Structural Cues
by: Ge, Fudong, et al.
Published: (2024)
by: Ge, Fudong, et al.
Published: (2024)
DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving
by: Shi, Chen, et al.
Published: (2025)
by: Shi, Chen, et al.
Published: (2025)
The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models
by: Mao, Runhao, et al.
Published: (2026)
by: Mao, Runhao, et al.
Published: (2026)
ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving
by: Peng, Qihang, et al.
Published: (2025)
by: Peng, Qihang, et al.
Published: (2025)
DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions
by: Zheng, Weicheng, et al.
Published: (2026)
by: Zheng, Weicheng, et al.
Published: (2026)
PFSD: A Multi-Modal Pedestrian-Focus Scene Dataset for Rich Tasks in Semi-Structured Environments
by: Liu, Yueting, et al.
Published: (2025)
by: Liu, Yueting, et al.
Published: (2025)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
by: Jiang, Hao, et al.
Published: (2025)
by: Jiang, Hao, et al.
Published: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
by: Zhang, Yiwei, et al.
Published: (2024)
by: Zhang, Yiwei, et al.
Published: (2024)
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous Driving
by: Chen, Xuesong, et al.
Published: (2025)
by: Chen, Xuesong, et al.
Published: (2025)
Spatial-aware Vision Language Model for Autonomous Driving
by: Wei, Weijie, et al.
Published: (2025)
by: Wei, Weijie, et al.
Published: (2025)
AutoPrune: Each Complexity Deserves a Pruning Policy
by: Wang, Hanshi, et al.
Published: (2025)
by: Wang, Hanshi, et al.
Published: (2025)
UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
by: Shi, Chen, et al.
Published: (2025)
by: Shi, Chen, et al.
Published: (2025)
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
by: Sun, Bin, et al.
Published: (2025)
by: Sun, Bin, et al.
Published: (2025)
DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
by: Chen, Yuntao, et al.
Published: (2024)
by: Chen, Yuntao, et al.
Published: (2024)
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
by: Chi, Haohan, et al.
Published: (2025)
by: Chi, Haohan, et al.
Published: (2025)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
Unifying Language-Action Understanding and Generation for Autonomous Driving
by: Wang, Xinyang, et al.
Published: (2026)
by: Wang, Xinyang, et al.
Published: (2026)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
by: Huang, Wenhui, et al.
Published: (2026)
by: Huang, Wenhui, et al.
Published: (2026)
Visual Adversarial Attack on Vision-Language Models for Autonomous Driving
by: Zhang, Tianyuan, et al.
Published: (2024)
by: Zhang, Tianyuan, et al.
Published: (2024)
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
by: Yang, Lijin, et al.
Published: (2026)
by: Yang, Lijin, et al.
Published: (2026)
ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
by: Li, Jingyu, et al.
Published: (2025)
by: Li, Jingyu, et al.
Published: (2025)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
SoLA-Vision: Fine-grained Layer-wise Linear Softmax Hybrid Attention
by: Li, Ruibang, et al.
Published: (2026)
by: Li, Ruibang, et al.
Published: (2026)
Learning Vision-Language-Action World Models for Autonomous Driving
by: Wang, Guoqing, et al.
Published: (2026)
by: Wang, Guoqing, et al.
Published: (2026)
VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving
by: Zhang, Haiming, et al.
Published: (2024)
by: Zhang, Haiming, et al.
Published: (2024)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026)
by: Mei, Xiaodong, et al.
Published: (2026)
Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving
by: Chen, Xuyang, et al.
Published: (2026)
by: Chen, Xuyang, et al.
Published: (2026)
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
by: Zhang, Jiaru, et al.
Published: (2026)
by: Zhang, Jiaru, et al.
Published: (2026)
A Unified Perception-Language-Action Framework for Adaptive Autonomous Driving
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving
by: Zhang, Enming, et al.
Published: (2024)
by: Zhang, Enming, et al.
Published: (2024)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
by: Xiong, Minhao, et al.
Published: (2025)
by: Xiong, Minhao, et al.
Published: (2025)
Similar Items
-
Integrating Diverse Assignment Strategies into DETRs
by: Zhang, Yiwei, et al.
Published: (2026) -
MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection
by: Wang, Hanshi, et al.
Published: (2025) -
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
by: Chen, Xuesong, et al.
Published: (2025) -
Online Segment Any 3D Thing as Instance Tracking
by: Wang, Hanshi, et al.
Published: (2025) -
DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving
by: Shi, Chen, et al.
Published: (2026)