DriveVA: Video Action Models are Zero-Shot Drivers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Mengmeng, Zhang, Diankun, Liu, Jiuming, Cui, Jianfeng, Xie, Hongwei, Chen, Guang, Ye, Hangjun, Yang, Michael Ying, Nex, Francesco, Cheng, Hao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915917493960704
author Liu, Mengmeng
Zhang, Diankun
Liu, Jiuming
Cui, Jianfeng
Xie, Hongwei
Chen, Guang
Ye, Hangjun
Yang, Michael Ying
Nex, Francesco
Cheng, Hao
author_facet Liu, Mengmeng
Zhang, Diankun
Liu, Jiuming
Cui, Jianfeng
Xie, Hongwei
Chen, Guang
Ye, Hangjun
Yang, Michael Ying
Nex, Francesco
Cheng, Hao
contents Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditions. Recent world-model-based planning methods have shown strong capabilities in scene understanding and multi-modal future prediction, yet their generalization across datasets and sensor configurations remains limited. In addition, their loosely coupled planning paradigm often leads to poor video-trajectory consistency during visual imagination. To overcome these limitations, we propose DriveVA, a novel autonomous driving world model that jointly decodes future visual forecasts and action sequences in a shared latent generative process. DriveVA inherits rich priors on motion dynamics and physical plausibility from well-pretrained large-scale video generation models to capture continuous spatiotemporal evolution and causal interaction patterns. To this end, DriveVA employs a DiT-based decoder to jointly predict future action sequences (trajectories) and videos, enabling tighter alignment between planning and scene evolution. We also introduce a video continuation strategy to strengthen long-duration rollout consistency. DriveVA achieves an impressive closed-loop performance of 90.9 PDM score on the challenge NAVSIM. Extensive experiments also demonstrate the zero-shot capability and cross-domain generalization of DriveVA, which reduces average L2 error and collision rate by 78.9% and 83.3% on nuScenes and 52.5% and 52.4% on the Bench2drive built on CARLA v2 compared with the state-of-the-art world-model-based planner.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04198
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DriveVA: Video Action Models are Zero-Shot Drivers
Liu, Mengmeng
Zhang, Diankun
Liu, Jiuming
Cui, Jianfeng
Xie, Hongwei
Chen, Guang
Ye, Hangjun
Yang, Michael Ying
Nex, Francesco
Cheng, Hao
Computer Vision and Pattern Recognition
Robotics
Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditions. Recent world-model-based planning methods have shown strong capabilities in scene understanding and multi-modal future prediction, yet their generalization across datasets and sensor configurations remains limited. In addition, their loosely coupled planning paradigm often leads to poor video-trajectory consistency during visual imagination. To overcome these limitations, we propose DriveVA, a novel autonomous driving world model that jointly decodes future visual forecasts and action sequences in a shared latent generative process. DriveVA inherits rich priors on motion dynamics and physical plausibility from well-pretrained large-scale video generation models to capture continuous spatiotemporal evolution and causal interaction patterns. To this end, DriveVA employs a DiT-based decoder to jointly predict future action sequences (trajectories) and videos, enabling tighter alignment between planning and scene evolution. We also introduce a video continuation strategy to strengthen long-duration rollout consistency. DriveVA achieves an impressive closed-loop performance of 90.9 PDM score on the challenge NAVSIM. Extensive experiments also demonstrate the zero-shot capability and cross-domain generalization of DriveVA, which reduces average L2 error and collision rate by 78.9% and 83.3% on nuScenes and 52.5% and 52.4% on the Bench2drive built on CARLA v2 compared with the state-of-the-art world-model-based planner.
title DriveVA: Video Action Models are Zero-Shot Drivers
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2604.04198