DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yuntao, Wang, Yuqi, Zhang, Zhaoxiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916541805625344
author Chen, Yuntao
Wang, Yuqi
Zhang, Zhaoxiang
author_facet Chen, Yuntao
Wang, Yuqi
Zhang, Zhaoxiang
contents World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual generation but lack the flexibility to incorporate other modalities like action. In contrast, autoregressive transformers have demonstrated exceptional capability in modeling multimodal data. Our work aims to unify both driving model simulation and trajectory planning into a single sequence modeling problem. We introduce a multimodal driving language based on interleaved image and action tokens, and develop DrivingGPT to learn joint world modeling and planning through standard next-token prediction. Our DrivingGPT demonstrates strong performance in both action-conditioned video generation and end-to-end planning, outperforming strong baselines on large-scale nuPlan and NAVSIM benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18607
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
Chen, Yuntao
Wang, Yuqi
Zhang, Zhaoxiang
Computer Vision and Pattern Recognition
World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual generation but lack the flexibility to incorporate other modalities like action. In contrast, autoregressive transformers have demonstrated exceptional capability in modeling multimodal data. Our work aims to unify both driving model simulation and trajectory planning into a single sequence modeling problem. We introduce a multimodal driving language based on interleaved image and action tokens, and develop DrivingGPT to learn joint world modeling and planning through standard next-token prediction. Our DrivingGPT demonstrates strong performance in both action-conditioned video generation and end-to-end planning, outperforming strong baselines on large-scale nuPlan and NAVSIM benchmarks.
title DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.18607