GEM: Generating LiDAR World Model via Deformable Mamba

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yang, Liu, Zhaojiang, Meng, Qiang, Liu, Youquan, Weng, Renliang, Qian, Jianjun, Yang, Jian, Xie, Jin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911661819953152
author Wu, Yang
Liu, Zhaojiang
Meng, Qiang
Liu, Youquan
Weng, Renliang
Qian, Jianjun
Yang, Jian
Xie, Jin
author_facet Wu, Yang
Liu, Zhaojiang
Meng, Qiang
Liu, Youquan
Weng, Renliang
Qian, Jianjun
Yang, Jian
Xie, Jin
contents World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, progress in LiDAR-based world models has lagged behind those built on camera videos or occupancy data, primarily due to two core challenges: the inherent disorder of LiDAR point clouds and the difficulty of distinguishing dynamic objects from static structures. To address these issues, we propose GEM: a Generative LiDAR world model that leverages deformable mamba architecture, significantly improving fidelity and imaginative capability. Specifically, leveraging the structural similarity between sequential laser scanning and Mamba's processing mechanism, we first tokenize LiDAR sweeps into compact representations via a custom LiDAR scene tokenizer. After unsupervised disentanglement of tokenized features via a dynamic-static separator, a tri-path deformable Mamba is introduced to perform selective scanning and adaptive gating fusion over the disentangled features, leading to enhanced spatial-temporal understanding of the world evolution. Optionally, a planner and a BEV layout controller can be integrated to explore the model's capability for autonomous rollout and its potential to generate ``what-if" scenarios. Extensive experiments show that GEM achieves state-of-the-art performances across diverse benchmarks and evaluation settings, demonstrating its superiority and effectiveness. Project page: https://github.com/wuyang98/GEM.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07326
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GEM: Generating LiDAR World Model via Deformable Mamba
Wu, Yang
Liu, Zhaojiang
Meng, Qiang
Liu, Youquan
Weng, Renliang
Qian, Jianjun
Yang, Jian
Xie, Jin
Computer Vision and Pattern Recognition
World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, progress in LiDAR-based world models has lagged behind those built on camera videos or occupancy data, primarily due to two core challenges: the inherent disorder of LiDAR point clouds and the difficulty of distinguishing dynamic objects from static structures. To address these issues, we propose GEM: a Generative LiDAR world model that leverages deformable mamba architecture, significantly improving fidelity and imaginative capability. Specifically, leveraging the structural similarity between sequential laser scanning and Mamba's processing mechanism, we first tokenize LiDAR sweeps into compact representations via a custom LiDAR scene tokenizer. After unsupervised disentanglement of tokenized features via a dynamic-static separator, a tri-path deformable Mamba is introduced to perform selective scanning and adaptive gating fusion over the disentangled features, leading to enhanced spatial-temporal understanding of the world evolution. Optionally, a planner and a BEV layout controller can be integrated to explore the model's capability for autonomous rollout and its potential to generate ``what-if" scenarios. Extensive experiments show that GEM achieves state-of-the-art performances across diverse benchmarks and evaluation settings, demonstrating its superiority and effectiveness. Project page: https://github.com/wuyang98/GEM.
title GEM: Generating LiDAR World Model via Deformable Mamba
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.07326