Saved in:
Bibliographic Details
Main Authors: Zhu, Dekai, Hu, Yixuan, Liu, Youquan, Lu, Dongyue, Kong, Lingdong, Ilic, Slobodan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.22643
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909935920480256
author Zhu, Dekai
Hu, Yixuan
Liu, Youquan
Lu, Dongyue
Kong, Lingdong
Ilic, Slobodan
author_facet Zhu, Dekai
Hu, Yixuan
Liu, Youquan
Lu, Dongyue
Kong, Lingdong
Ilic, Slobodan
contents Leveraging recent diffusion models, LiDAR-based large-scale 3D scene generation has achieved great success. While recent voxel-based approaches can generate both geometric structures and semantic labels, existing range-view methods are limited to producing unlabeled LiDAR scenes. Relying on pretrained segmentation models to predict the semantic maps often results in suboptimal cross-modal consistency. To address this limitation while preserving the advantages of range-view representations, such as computational efficiency and simplified network design, we propose Spiral, a novel range-view LiDAR diffusion model that simultaneously generates depth, reflectance images, and semantic maps. Furthermore, we introduce novel semantic-aware metrics to evaluate the quality of the generated labeled range-view data. Experiments on the SemanticKITTI and nuScenes datasets demonstrate that Spiral achieves state-of-the-art performance with the smallest parameter size, outperforming two-step methods that combine the generative and segmentation models. Additionally, we validate that range images generated by Spiral can be effectively used for synthetic data augmentation in the downstream segmentation training, significantly reducing the labeling effort on LiDAR data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22643
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPIRAL: Semantic-Aware Progressive LiDAR Scene Generation and Understanding
Zhu, Dekai
Hu, Yixuan
Liu, Youquan
Lu, Dongyue
Kong, Lingdong
Ilic, Slobodan
Computer Vision and Pattern Recognition
Leveraging recent diffusion models, LiDAR-based large-scale 3D scene generation has achieved great success. While recent voxel-based approaches can generate both geometric structures and semantic labels, existing range-view methods are limited to producing unlabeled LiDAR scenes. Relying on pretrained segmentation models to predict the semantic maps often results in suboptimal cross-modal consistency. To address this limitation while preserving the advantages of range-view representations, such as computational efficiency and simplified network design, we propose Spiral, a novel range-view LiDAR diffusion model that simultaneously generates depth, reflectance images, and semantic maps. Furthermore, we introduce novel semantic-aware metrics to evaluate the quality of the generated labeled range-view data. Experiments on the SemanticKITTI and nuScenes datasets demonstrate that Spiral achieves state-of-the-art performance with the smallest parameter size, outperforming two-step methods that combine the generative and segmentation models. Additionally, we validate that range images generated by Spiral can be effectively used for synthetic data augmentation in the downstream segmentation training, significantly reducing the labeling effort on LiDAR data.
title SPIRAL: Semantic-Aware Progressive LiDAR Scene Generation and Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22643