DriveVA: Video Action Models are Zero-Shot Drivers
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Mengmeng, Zhang, Diankun, Liu, Jiuming, Cui, Jianfeng, Xie, Hongwei, Chen, Guang, Ye, Hangjun, Yang, Michael Ying, Nex, Francesco, Cheng, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026)
by: Mei, Xiaodong, et al.
Published: (2026)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
ZeD-MAP: Bundle Adjustment Guided Zero-Shot Depth Maps for Real-Time Aerial Imaging
by: Iz, Selim Ahmet, et al.
Published: (2026)
by: Iz, Selim Ahmet, et al.
Published: (2026)
4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation
by: Liu, Mengmeng, et al.
Published: (2025)
by: Liu, Mengmeng, et al.
Published: (2025)
SocialNav-Map: Dynamic Mapping with Human Trajectory Prediction for Zero-Shot Social Navigation
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
TopoLiDM: Topology-Aware LiDAR Diffusion Models for Interpretable and Realistic LiDAR Point Cloud Generation
by: Liu, Jiuming, et al.
Published: (2025)
by: Liu, Jiuming, et al.
Published: (2025)
Learning from Mistakes: Post-Training for Driving VLA with Takeover Data
by: Gao, Yinfeng, et al.
Published: (2026)
by: Gao, Yinfeng, et al.
Published: (2026)
NavDreamer: Video Models as Zero-Shot 3D Navigators
by: Huang, Xijie, et al.
Published: (2026)
by: Huang, Xijie, et al.
Published: (2026)
An Analysis of Driver-Initiated Takeovers during Assisted Driving and their Effect on Driver Satisfaction
by: Schwager, Robin, et al.
Published: (2024)
by: Schwager, Robin, et al.
Published: (2024)
EmboAlign: Aligning Video Generation with Compositional Constraints for Zero-Shot Manipulation
by: Zhang, Gehao, et al.
Published: (2026)
by: Zhang, Gehao, et al.
Published: (2026)
Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving
by: Zheng, Yinan, et al.
Published: (2026)
by: Zheng, Yinan, et al.
Published: (2026)
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
by: Wang, Junli, et al.
Published: (2026)
by: Wang, Junli, et al.
Published: (2026)
PerlAD: Towards Enhanced Closed-loop End-to-end Autonomous Driving with Pseudo-simulation-based Reinforcement Learning
by: Gao, Yinfeng, et al.
Published: (2026)
by: Gao, Yinfeng, et al.
Published: (2026)
MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving
by: Wang, Junli, et al.
Published: (2026)
by: Wang, Junli, et al.
Published: (2026)
Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping
by: Udugama, U. V. B. L., et al.
Published: (2026)
by: Udugama, U. V. B. L., et al.
Published: (2026)
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
Jointly Learning Predicates and Actions Enables Zero-Shot Skill Composition
by: Quartey, Benedict, et al.
Published: (2026)
by: Quartey, Benedict, et al.
Published: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
Biasing the Driving Style of an Artificial Race Driver for Online Time-Optimal Maneuver Planning
by: Taddei, Sebastiano, et al.
Published: (2025)
by: Taddei, Sebastiano, et al.
Published: (2025)
UniManip: General-Purpose Zero-Shot Robotic Manipulation with Agentic Operational Graph
by: Liu, Haichao, et al.
Published: (2026)
by: Liu, Haichao, et al.
Published: (2026)
Generalizing End-To-End Autonomous Driving In Real-World Environments Using Zero-Shot LLMs
by: Dong, Zeyu, et al.
Published: (2024)
by: Dong, Zeyu, et al.
Published: (2024)
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
by: Wen, Congcong, et al.
Published: (2025)
by: Wen, Congcong, et al.
Published: (2025)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
by: Chen, Guangyan, et al.
Published: (2025)
by: Chen, Guangyan, et al.
Published: (2025)
MVAdapt: Zero-Shot Multi-Vehicle Adaptation for End-to-End Autonomous Driving
by: Oh, Haesung, et al.
Published: (2026)
by: Oh, Haesung, et al.
Published: (2026)
M2H: Multi-Task Learning with Efficient Window-Based Cross-Task Attention for Monocular Spatial Perception
by: Udugama, U. V. B. L, et al.
Published: (2025)
by: Udugama, U. V. B. L, et al.
Published: (2025)
Indirect Shared Control of Highly Automated Vehicles for Cooperative Driving between Driver and Automation
by: Li, Renjie, et al.
Published: (2017)
by: Li, Renjie, et al.
Published: (2017)
SimScale: Learning to Drive via Real-World Simulation at Scale
by: Tian, Haochen, et al.
Published: (2025)
by: Tian, Haochen, et al.
Published: (2025)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
by: Liu, Jinkun, et al.
Published: (2026)
by: Liu, Jinkun, et al.
Published: (2026)
Multi-Floor Zero-Shot Object Navigation Policy
by: Zhang, Lingfeng, et al.
Published: (2024)
by: Zhang, Lingfeng, et al.
Published: (2024)
AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement
by: Hu, Zhaofeng, et al.
Published: (2026)
by: Hu, Zhaofeng, et al.
Published: (2026)
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models
by: Heo, Hyeongjun, et al.
Published: (2026)
by: Heo, Hyeongjun, et al.
Published: (2026)
Leveraging Semantic and Geometric Information for Zero-Shot Robot-to-Human Handover
by: Liu, Jiangshan, et al.
Published: (2024)
by: Liu, Jiangshan, et al.
Published: (2024)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
by: Yuan, Shuaihang, et al.
Published: (2024)
by: Yuan, Shuaihang, et al.
Published: (2024)
TriHelper: Zero-Shot Object Navigation with Dynamic Assistance
by: Zhang, Lingfeng, et al.
Published: (2024)
by: Zhang, Lingfeng, et al.
Published: (2024)
USS-Nav: Unified Spatio-Semantic Scene Graph for Lightweight UAV Zero-Shot Object Navigation
by: Gai, Weiqi, et al.
Published: (2026)
by: Gai, Weiqi, et al.
Published: (2026)
Zero-Shot Adaptation to Robot Structural Damage via Natural Language-Informed Kinodynamics Modeling
by: Pokhrel, Anuj, et al.
Published: (2026)
by: Pokhrel, Anuj, et al.
Published: (2026)
Is Your VLM for Autonomous Driving Safety-Ready? A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
by: Meng, Xianhui, et al.
Published: (2025)
by: Meng, Xianhui, et al.
Published: (2025)
Similar Items
-
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025) -
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026) -
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026) -
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026) -
ZeD-MAP: Bundle Adjustment Guided Zero-Shot Depth Maps for Real-Time Aerial Imaging
by: Iz, Selim Ahmet, et al.
Published: (2026)