GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Tianchen, Chen, Xuefeng, Chen, Yi, Chen, Qu, Xu, Yuyao, Yang, Lijin, Xu, Le, Zhang, Yu, Zhang, Bo, Huang, Wuxiong, Wang, Hesheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909048072306688
author Deng, Tianchen
Chen, Xuefeng
Chen, Yi
Chen, Qu
Xu, Yuyao
Yang, Lijin
Xu, Le
Zhang, Yu
Zhang, Bo
Huang, Wuxiong
Wang, Hesheng
author_facet Deng, Tianchen
Chen, Xuefeng
Chen, Yi
Chen, Qu
Xu, Yuyao
Yang, Lijin
Xu, Le
Zhang, Yu
Zhang, Bo
Huang, Wuxiong
Wang, Hesheng
contents Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to interpret or reason about the driving environment. Moreover, current approaches represent 3D spatial information with point cloud or BEV features do not accurately align textual information with the underlying 3D scene. To address these limitations, we propose a novel unified DWM framework based on 3D Gaussian scene representation, which enables both 3D scene understanding and multi-modal scene generation, while also enabling contextual enrichment for understanding and generation tasks. Our approach directly aligns textual information with the 3D scene by embedding rich linguistic features into each Gaussian primitive, thereby achieving early modality alignment. In addition, we design a novel task-aware language-guided sampling strategy that removes redundant 3D Gaussians and injects accurate and compact 3D tokens into LLM. Furthermore, we design a dual-condition multi-modal generation model, where the information captured by our vision-language model is leveraged as a high-level language condition in combination with a low-level image condition, jointly guiding the multi-modal generation process. We conduct comprehensive studies on the nuScenes, and NuInteract datasets to validate the effectiveness of our framework. Our method achieves state-of-the-art performance. We will release the code publicly on GitHub https://github.com/dtc111111/GaussianDWM.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23180
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
Deng, Tianchen
Chen, Xuefeng
Chen, Yi
Chen, Qu
Xu, Yuyao
Yang, Lijin
Xu, Le
Zhang, Yu
Zhang, Bo
Huang, Wuxiong
Wang, Hesheng
Computer Vision and Pattern Recognition
Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to interpret or reason about the driving environment. Moreover, current approaches represent 3D spatial information with point cloud or BEV features do not accurately align textual information with the underlying 3D scene. To address these limitations, we propose a novel unified DWM framework based on 3D Gaussian scene representation, which enables both 3D scene understanding and multi-modal scene generation, while also enabling contextual enrichment for understanding and generation tasks. Our approach directly aligns textual information with the 3D scene by embedding rich linguistic features into each Gaussian primitive, thereby achieving early modality alignment. In addition, we design a novel task-aware language-guided sampling strategy that removes redundant 3D Gaussians and injects accurate and compact 3D tokens into LLM. Furthermore, we design a dual-condition multi-modal generation model, where the information captured by our vision-language model is leveraged as a high-level language condition in combination with a low-level image condition, jointly guiding the multi-modal generation process. We conduct comprehensive studies on the nuScenes, and NuInteract datasets to validate the effectiveness of our framework. Our method achieves state-of-the-art performance. We will release the code publicly on GitHub https://github.com/dtc111111/GaussianDWM.
title GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.23180