LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Pei, Wang, Songtao, Zhang, Lang, Peng, Xingyue, Lyu, Yuandong, Deng, Jiaxin, Lu, Songxin, Ma, Weiliang, Zhang, Xueyang, Zhan, Yifei, Lang, XianPeng, Ma, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912720378396672
author Liu, Pei
Wang, Songtao
Zhang, Lang
Peng, Xingyue
Lyu, Yuandong
Deng, Jiaxin
Lu, Songxin
Ma, Weiliang
Zhang, Xueyang
Zhan, Yifei
Lang, XianPeng
Ma, Jun
author_facet Liu, Pei
Wang, Songtao
Zhang, Lang
Peng, Xingyue
Lyu, Yuandong
Deng, Jiaxin
Lu, Songxin
Ma, Weiliang
Zhang, Xueyang
Zhan, Yifei
Lang, XianPeng
Ma, Jun
contents Synthesizing high-fidelity and controllable 4D LiDAR data is crucial for creating scalable simulation environments for autonomous driving. This task is inherently challenging due to the sensor's unique spherical geometry, the temporal sparsity of point clouds, and the complexity of dynamic scenes. To address these challenges, we present LiSTAR, a novel generative world model that operates directly on the sensor's native geometry. LiSTAR introduces a Hybrid-Cylindrical-Spherical (HCS) representation to preserve data fidelity by mitigating quantization artifacts common in Cartesian grids. To capture complex dynamics from sparse temporal data, it utilizes a Spatio-Temporal Attention with Ray-Centric Transformer (START) that explicitly models feature evolution along individual sensor rays for robust temporal coherence. Furthermore, for controllable synthesis, we propose a novel 4D point cloud-aligned voxel layout for conditioning and a corresponding discrete Masked Generative START (MaskSTART) framework, which learns a compact, tokenized representation of the scene, enabling efficient, high-resolution, and layout-guided compositional generation. Comprehensive experiments validate LiSTAR's state-of-the-art performance across 4D LiDAR reconstruction, prediction, and conditional generation, with substantial quantitative gains: reducing generation MMD by a massive 76%, improving reconstruction IoU by 32%, and lowering prediction L1 Med by 50%. This level of performance provides a powerful new foundation for creating realistic and controllable autonomous systems simulations. Project link: https://ocean-luna.github.io/LiSTAR.gitub.io.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
Liu, Pei
Wang, Songtao
Zhang, Lang
Peng, Xingyue
Lyu, Yuandong
Deng, Jiaxin
Lu, Songxin
Ma, Weiliang
Zhang, Xueyang
Zhan, Yifei
Lang, XianPeng
Ma, Jun
Computer Vision and Pattern Recognition
Synthesizing high-fidelity and controllable 4D LiDAR data is crucial for creating scalable simulation environments for autonomous driving. This task is inherently challenging due to the sensor's unique spherical geometry, the temporal sparsity of point clouds, and the complexity of dynamic scenes. To address these challenges, we present LiSTAR, a novel generative world model that operates directly on the sensor's native geometry. LiSTAR introduces a Hybrid-Cylindrical-Spherical (HCS) representation to preserve data fidelity by mitigating quantization artifacts common in Cartesian grids. To capture complex dynamics from sparse temporal data, it utilizes a Spatio-Temporal Attention with Ray-Centric Transformer (START) that explicitly models feature evolution along individual sensor rays for robust temporal coherence. Furthermore, for controllable synthesis, we propose a novel 4D point cloud-aligned voxel layout for conditioning and a corresponding discrete Masked Generative START (MaskSTART) framework, which learns a compact, tokenized representation of the scene, enabling efficient, high-resolution, and layout-guided compositional generation. Comprehensive experiments validate LiSTAR's state-of-the-art performance across 4D LiDAR reconstruction, prediction, and conditional generation, with substantial quantitative gains: reducing generation MMD by a massive 76%, improving reconstruction IoU by 32%, and lowering prediction L1 Med by 50%. This level of performance provides a powerful new foundation for creating realistic and controllable autonomous systems simulations. Project link: https://ocean-luna.github.io/LiSTAR.gitub.io.
title LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.16049