SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ruoyu, Wang, Jingke, Ma, Yukai, Huang, Yuehao, Lei, Shuangming, Xu, Guanglin, Ye, Aixue, Liu, Yong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916041493315584
author Wang, Ruoyu
Wang, Jingke
Ma, Yukai
Huang, Yuehao
Lei, Shuangming
Xu, Guanglin
Ye, Aixue
Liu, Yong
author_facet Wang, Ruoyu
Wang, Jingke
Ma, Yukai
Huang, Yuehao
Lei, Shuangming
Xu, Guanglin
Ye, Aixue
Liu, Yong
contents Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, existing driving world models are typically built upon dense scene representations, causing high computational costs and redundant information. In this paper, we present SparseWorld, a lightweight world model that focuses on predicting only the critical layout of the scene, enabling efficient future forecasting for end-to-end driving systems. SparseWorld first performs autoregressive rollout to forecast future map elements and surrounding agents, enabling the model to learn how driving scenarios evolve over time. It then leverages these predicted futures to refine downstream motion prediction and trajectory planning. Specifically, we propose a Sparse Dreamer that anticipates future instances in the latent space through joint temporal and spatial attention. By interacting with predicted future instances, the motion planner captures more accurate motion patterns and generates more informed and safety-aware trajectories. Extensive experiments demonstrate that SparseWorld significantly reduces collision risk and achieves state-of-the-art performance on the open-loop planning metrics of the nuScenes dataset with a collision rate of 0.05\%. Moreover, it substantially outperforms the baseline method in closed-loop planning metrics on the Bench2Drive benchmark. Supplementary material is available at the project page: https://wryzju.github.io/SparseWorld/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24354
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
Wang, Ruoyu
Wang, Jingke
Ma, Yukai
Huang, Yuehao
Lei, Shuangming
Xu, Guanglin
Ye, Aixue
Liu, Yong
Computer Vision and Pattern Recognition
Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, existing driving world models are typically built upon dense scene representations, causing high computational costs and redundant information. In this paper, we present SparseWorld, a lightweight world model that focuses on predicting only the critical layout of the scene, enabling efficient future forecasting for end-to-end driving systems. SparseWorld first performs autoregressive rollout to forecast future map elements and surrounding agents, enabling the model to learn how driving scenarios evolve over time. It then leverages these predicted futures to refine downstream motion prediction and trajectory planning. Specifically, we propose a Sparse Dreamer that anticipates future instances in the latent space through joint temporal and spatial attention. By interacting with predicted future instances, the motion planner captures more accurate motion patterns and generates more informed and safety-aware trajectories. Extensive experiments demonstrate that SparseWorld significantly reduces collision risk and achieves state-of-the-art performance on the open-loop planning metrics of the nuScenes dataset with a collision rate of 0.05\%. Moreover, it substantially outperforms the baseline method in closed-loop planning metrics on the Bench2Drive benchmark. Supplementary material is available at the project page: https://wryzju.github.io/SparseWorld/.
title SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.24354