DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917414985269248 |
|---|---|
| author | Xia, Tianze Li, Yongkang Zhou, Lijun Yao, Jingfeng Xiong, Kaixin Sun, Haiyang Wang, Bing Ma, Kun Chen, Guang Ye, Hangjun Liu, Wenyu Wang, Xinggang |
| author_facet | Xia, Tianze Li, Yongkang Zhou, Lijun Yao, Jingfeng Xiong, Kaixin Sun, Haiyang Wang, Bing Ma, Kun Chen, Guang Ye, Hangjun Liu, Wenyu Wang, Xinggang |
| contents | World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate within ostensibly unified architectures that still keep world prediction and motion planning as decoupled processes. To bridge this gap, we propose DriveLaW, a novel paradigm that unifies video generation and motion planning. By directly injecting the latent representation from its video generator into the planner, DriveLaW ensures inherent consistency between high-fidelity future generation and reliable trajectory planning. Specifically, DriveLaW consists of two core components: DriveLaW-Video, our powerful world model that generates high-fidelity forecasting with expressive latent representations, and DriveLaW-Act, a diffusion planner that generates consistent and reliable trajectories from the latent of DriveLaW-Video, with both components optimized by a three-stage progressive training strategy. The power of our unified paradigm is demonstrated by new state-of-the-art results across both tasks. DriveLaW not only advances video prediction significantly, surpassing best-performing work by 33.3% in FID and 1.8% in FVD, but also achieves a new record on the NAVSIM planning benchmark. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_23421 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DriveLaW:Unifying Planning and Video Generation in a Latent Driving World Xia, Tianze Li, Yongkang Zhou, Lijun Yao, Jingfeng Xiong, Kaixin Sun, Haiyang Wang, Bing Ma, Kun Chen, Guang Ye, Hangjun Liu, Wenyu Wang, Xinggang Computer Vision and Pattern Recognition World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate within ostensibly unified architectures that still keep world prediction and motion planning as decoupled processes. To bridge this gap, we propose DriveLaW, a novel paradigm that unifies video generation and motion planning. By directly injecting the latent representation from its video generator into the planner, DriveLaW ensures inherent consistency between high-fidelity future generation and reliable trajectory planning. Specifically, DriveLaW consists of two core components: DriveLaW-Video, our powerful world model that generates high-fidelity forecasting with expressive latent representations, and DriveLaW-Act, a diffusion planner that generates consistent and reliable trajectories from the latent of DriveLaW-Video, with both components optimized by a three-stage progressive training strategy. The power of our unified paradigm is demonstrated by new state-of-the-art results across both tasks. DriveLaW not only advances video prediction significantly, surpassing best-performing work by 33.3% in FID and 1.8% in FVD, but also achieves a new record on the NAVSIM planning benchmark. |
| title | DriveLaW:Unifying Planning and Video Generation in a Latent Driving World |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.23421 |