SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Shuhan, Lambert, John, Jeon, Hong, Kulshrestha, Sakshum, Bai, Yijing, Luo, Jing, Anguelov, Dragomir, Tan, Mingxing, Jiang, Chiyu Max
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909662128898048
author Tan, Shuhan
Lambert, John
Jeon, Hong
Kulshrestha, Sakshum
Bai, Yijing
Luo, Jing
Anguelov, Dragomir
Tan, Mingxing
Jiang, Chiyu Max
author_facet Tan, Shuhan
Lambert, John
Jeon, Hong
Kulshrestha, Sakshum
Bai, Yijing
Luo, Jing
Anguelov, Dragomir
Tan, Mingxing
Jiang, Chiyu Max
contents The goal of traffic simulation is to augment a potentially limited amount of manually-driven miles that is available for testing and validation, with a much larger amount of simulated synthetic miles. The culmination of this vision would be a generative simulated city, where given a map of the city and an autonomous vehicle (AV) software stack, the simulator can seamlessly simulate the trip from point A to point B by populating the city around the AV and controlling all aspects of the scene, from animating the dynamic agents (e.g., vehicles, pedestrians) to controlling the traffic light states. We refer to this vision as CitySim, which requires an agglomeration of simulation technologies: scene generation to populate the initial scene, agent behavior modeling to animate the scene, occlusion reasoning, dynamic scene generation to seamlessly spawn and remove agents, and environment simulation for factors such as traffic lights. While some key technologies have been separately studied in various works, others such as dynamic scene generation and environment simulation have received less attention in the research community. We propose SceneDiffuser++, the first end-to-end generative world model trained on a single loss function capable of point A-to-B simulation on a city scale integrating all the requirements above. We demonstrate the city-scale traffic simulation capability of SceneDiffuser++ and study its superior realism under long simulation conditions. We evaluate the simulation quality on an augmented version of the Waymo Open Motion Dataset (WOMD) with larger map regions to support trip-level simulation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21976
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model
Tan, Shuhan
Lambert, John
Jeon, Hong
Kulshrestha, Sakshum
Bai, Yijing
Luo, Jing
Anguelov, Dragomir
Tan, Mingxing
Jiang, Chiyu Max
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Multiagent Systems
Robotics
The goal of traffic simulation is to augment a potentially limited amount of manually-driven miles that is available for testing and validation, with a much larger amount of simulated synthetic miles. The culmination of this vision would be a generative simulated city, where given a map of the city and an autonomous vehicle (AV) software stack, the simulator can seamlessly simulate the trip from point A to point B by populating the city around the AV and controlling all aspects of the scene, from animating the dynamic agents (e.g., vehicles, pedestrians) to controlling the traffic light states. We refer to this vision as CitySim, which requires an agglomeration of simulation technologies: scene generation to populate the initial scene, agent behavior modeling to animate the scene, occlusion reasoning, dynamic scene generation to seamlessly spawn and remove agents, and environment simulation for factors such as traffic lights. While some key technologies have been separately studied in various works, others such as dynamic scene generation and environment simulation have received less attention in the research community. We propose SceneDiffuser++, the first end-to-end generative world model trained on a single loss function capable of point A-to-B simulation on a city scale integrating all the requirements above. We demonstrate the city-scale traffic simulation capability of SceneDiffuser++ and study its superior realism under long simulation conditions. We evaluate the simulation quality on an augmented version of the Waymo Open Motion Dataset (WOMD) with larger map regions to support trip-level simulation.
title SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Multiagent Systems
Robotics
url https://arxiv.org/abs/2506.21976