DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Xiaotao, Yin, Wei, Jia, Mingkai, Deng, Junyuan, Guo, Xiaoyang, Zhang, Qian, Long, Xiaoxiao, Tan, Ping
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913628609839104
author Hu, Xiaotao
Yin, Wei
Jia, Mingkai
Deng, Junyuan
Guo, Xiaoyang
Zhang, Qian
Long, Xiaoxiao
Tan, Ping
author_facet Hu, Xiaotao
Yin, Wei
Jia, Mingkai
Deng, Junyuan
Guo, Xiaoyang
Zhang, Qian
Long, Xiaoxiao
Tan, Ping
contents Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. Some works attempt to extend this approach to autonomous driving by building video-based world models capable of generating realistic future video sequences and predicting ego states. However, prior works tend to produce unsatisfactory results, as the classic GPT framework is designed to handle 1D contextual information, such as text, and lacks the inherent ability to model the spatial and temporal dynamics essential for video generation. In this paper, we present DrivingWorld, a GPT-style world model for autonomous driving, featuring several spatial-temporal fusion mechanisms. This design enables effective modeling of both spatial and temporal dynamics, facilitating high-fidelity, long-duration video generation. Specifically, we propose a next-state prediction strategy to model temporal coherence between consecutive frames and apply a next-token prediction strategy to capture spatial information within each frame. To further enhance generalization ability, we propose a novel masking strategy and reweighting strategy for token prediction to mitigate long-term drifting issues and enable precise control. Our work demonstrates the ability to produce high-fidelity and consistent video clips of over 40 seconds in duration, which is over 2 times longer than state-of-the-art driving world models. Experiments show that, in contrast to prior works, our method achieves superior visual quality and significantly more accurate controllable future video generation. Our code is available at https://github.com/YvanYin/DrivingWorld.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19505
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
Hu, Xiaotao
Yin, Wei
Jia, Mingkai
Deng, Junyuan
Guo, Xiaoyang
Zhang, Qian
Long, Xiaoxiao
Tan, Ping
Computer Vision and Pattern Recognition
Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. Some works attempt to extend this approach to autonomous driving by building video-based world models capable of generating realistic future video sequences and predicting ego states. However, prior works tend to produce unsatisfactory results, as the classic GPT framework is designed to handle 1D contextual information, such as text, and lacks the inherent ability to model the spatial and temporal dynamics essential for video generation. In this paper, we present DrivingWorld, a GPT-style world model for autonomous driving, featuring several spatial-temporal fusion mechanisms. This design enables effective modeling of both spatial and temporal dynamics, facilitating high-fidelity, long-duration video generation. Specifically, we propose a next-state prediction strategy to model temporal coherence between consecutive frames and apply a next-token prediction strategy to capture spatial information within each frame. To further enhance generalization ability, we propose a novel masking strategy and reweighting strategy for token prediction to mitigate long-term drifting issues and enable precise control. Our work demonstrates the ability to produce high-fidelity and consistent video clips of over 40 seconds in duration, which is over 2 times longer than state-of-the-art driving world models. Experiments show that, in contrast to prior works, our method achieves superior visual quality and significantly more accurate controllable future video generation. Our code is available at https://github.com/YvanYin/DrivingWorld.
title DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19505