WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhu, Ziyue, Wu, Zhanqian, Zhu, Zhenxin, Zhou, Lijun, Sun, Haiyang, Wan, Bing, Ma, Kun, Chen, Guang, Ye, Hangjun, Xie, Jin, Yang, jian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911214923153408
author Zhu, Ziyue
Wu, Zhanqian
Zhu, Zhenxin
Zhou, Lijun
Sun, Haiyang
Wan, Bing
Ma, Kun
Chen, Guang
Ye, Hangjun
Xie, Jin
Yang, jian
author_facet Zhu, Ziyue
Wu, Zhanqian
Zhu, Zhenxin
Zhou, Lijun
Sun, Haiyang
Wan, Bing
Ma, Kun
Chen, Guang
Ye, Hangjun
Xie, Jin
Yang, jian
contents Recent advances in driving-scene generation and reconstruction have demonstrated significant potential for enhancing autonomous driving systems by producing scalable and controllable training data. Existing generation methods primarily focus on synthesizing diverse and high-fidelity driving videos; however, due to limited 3D consistency and sparse viewpoint coverage, they struggle to support convenient and high-quality novel-view synthesis (NVS). Conversely, recent 3D/4D reconstruction approaches have significantly improved NVS for real-world driving scenes, yet inherently lack generative capabilities. To overcome this dilemma between scene generation and reconstruction, we propose WorldSplat, a novel feed-forward framework for 4D driving-scene generation. Our approach effectively generates consistent multi-track videos through two key steps: (i) We introduce a 4D-aware latent diffusion model integrating multi-modal information to produce pixel-aligned 4D Gaussians in a feed-forward manner. (ii) Subsequently, we refine the novel view videos rendered from these Gaussians using a enhanced video diffusion model. Extensive experiments conducted on benchmark datasets demonstrate that WorldSplat effectively generates high-fidelity, temporally and spatially consistent multi-track novel view driving videos. Project: https://wm-research.github.io/worldsplat/
format Preprint
id arxiv_https___arxiv_org_abs_2509_23402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
Zhu, Ziyue
Wu, Zhanqian
Zhu, Zhenxin
Zhou, Lijun
Sun, Haiyang
Wan, Bing
Ma, Kun
Chen, Guang
Ye, Hangjun
Xie, Jin
Yang, jian
Computer Vision and Pattern Recognition
Recent advances in driving-scene generation and reconstruction have demonstrated significant potential for enhancing autonomous driving systems by producing scalable and controllable training data. Existing generation methods primarily focus on synthesizing diverse and high-fidelity driving videos; however, due to limited 3D consistency and sparse viewpoint coverage, they struggle to support convenient and high-quality novel-view synthesis (NVS). Conversely, recent 3D/4D reconstruction approaches have significantly improved NVS for real-world driving scenes, yet inherently lack generative capabilities. To overcome this dilemma between scene generation and reconstruction, we propose WorldSplat, a novel feed-forward framework for 4D driving-scene generation. Our approach effectively generates consistent multi-track videos through two key steps: (i) We introduce a 4D-aware latent diffusion model integrating multi-modal information to produce pixel-aligned 4D Gaussians in a feed-forward manner. (ii) Subsequently, we refine the novel view videos rendered from these Gaussians using a enhanced video diffusion model. Extensive experiments conducted on benchmark datasets demonstrate that WorldSplat effectively generates high-fidelity, temporally and spatially consistent multi-track novel view driving videos. Project: https://wm-research.github.io/worldsplat/
title WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.23402