Epona: Autoregressive Diffusion World Model for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kaiwen, Tang, Zhenyu, Hu, Xiaotao, Pan, Xingang, Guo, Xiaoyang, Liu, Yuan, Huang, Jingwei, Yuan, Li, Zhang, Qian, Long, Xiao-Xiao, Cao, Xun, Yin, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908428356550656
author Zhang, Kaiwen
Tang, Zhenyu
Hu, Xiaotao
Pan, Xingang
Guo, Xiaoyang
Liu, Yuan
Huang, Jingwei
Yuan, Li
Zhang, Qian
Long, Xiao-Xiao
Cao, Xun
Yin, Wei
author_facet Zhang, Kaiwen
Tang, Zhenyu
Hu, Xiaotao
Pan, Xingang
Guo, Xiaoyang
Liu, Yuan
Huang, Jingwei
Yuan, Li
Zhang, Qian
Long, Xiao-Xiao
Cao, Xun
Yin, Wei
contents Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is because conventional video diffusion models rely on global joint distribution modeling of fixed-length frame sequences rather than sequentially constructing localized distributions at each timestep. In this work, we propose Epona, an autoregressive diffusion world model that enables localized spatiotemporal distribution modeling through two key innovations: 1) Decoupled spatiotemporal factorization that separates temporal dynamics modeling from fine-grained future world generation, and 2) Modular trajectory and video prediction that seamlessly integrate motion planning with visual modeling in an end-to-end framework. Our architecture enables high-resolution, long-duration generation while introducing a novel chain-of-forward training strategy to address error accumulation in autoregressive loops. Experimental results demonstrate state-of-the-art performance with 7.4\% FVD improvement and minutes longer prediction duration compared to prior works. The learned world model further serves as a real-time motion planner, outperforming strong end-to-end planners on NAVSIM benchmarks. Code will be publicly available at \href{https://github.com/Kevin-thu/Epona/}{https://github.com/Kevin-thu/Epona/}.
format Preprint
id arxiv_https___arxiv_org_abs_2506_24113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Epona: Autoregressive Diffusion World Model for Autonomous Driving
Zhang, Kaiwen
Tang, Zhenyu
Hu, Xiaotao
Pan, Xingang
Guo, Xiaoyang
Liu, Yuan
Huang, Jingwei
Yuan, Li
Zhang, Qian
Long, Xiao-Xiao
Cao, Xun
Yin, Wei
Computer Vision and Pattern Recognition
Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is because conventional video diffusion models rely on global joint distribution modeling of fixed-length frame sequences rather than sequentially constructing localized distributions at each timestep. In this work, we propose Epona, an autoregressive diffusion world model that enables localized spatiotemporal distribution modeling through two key innovations: 1) Decoupled spatiotemporal factorization that separates temporal dynamics modeling from fine-grained future world generation, and 2) Modular trajectory and video prediction that seamlessly integrate motion planning with visual modeling in an end-to-end framework. Our architecture enables high-resolution, long-duration generation while introducing a novel chain-of-forward training strategy to address error accumulation in autoregressive loops. Experimental results demonstrate state-of-the-art performance with 7.4\% FVD improvement and minutes longer prediction duration compared to prior works. The learned world model further serves as a real-time motion planner, outperforming strong end-to-end planners on NAVSIM benchmarks. Code will be publicly available at \href{https://github.com/Kevin-thu/Epona/}{https://github.com/Kevin-thu/Epona/}.
title Epona: Autoregressive Diffusion World Model for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.24113