Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Shengyuan, Zhao, An, Yang, Ling, Li, Zejian, Meng, Chenye, Xu, Haoran, Chen, Tianrun, Wei, AnYang, GU, Perry Pengyun, Sun, Lingyun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915412026851328
author Zhang, Shengyuan
Zhao, An
Yang, Ling
Li, Zejian
Meng, Chenye
Xu, Haoran
Chen, Tianrun
Wei, AnYang
GU, Perry Pengyun
Sun, Lingyun
author_facet Zhang, Shengyuan
Zhao, An
Yang, Ling
Li, Zejian
Meng, Chenye
Xu, Haoran
Chen, Tianrun
Wei, AnYang
GU, Perry Pengyun
Sun, Lingyun
contents Diffusion models have been applied to 3D LiDAR scene completion due to their strong training stability and high completion quality. However, the slow sampling speed limits the practical application of diffusion-based scene completion models since autonomous vehicles require an efficient perception of surrounding environments. This paper proposes a novel distillation method tailored for 3D Li- DAR scene completion models, dubbed ScoreLiDAR, which achieves efficient yet high-quality scene completion. Score- LiDAR enables the distilled model to sample in significantly fewer steps after distillation. To improve completion quality, we also introduce a novel Structural Loss, which encourages the distilled model to capture the geometric structure of the 3D LiDAR scene. The loss contains a scene-wise term constraining the holistic structure and a point-wise term constraining the key landmark points and their relative configuration. Extensive experiments demonstrate that ScoreLiDAR significantly accelerates the completion time from 30.55 to 5.37 seconds per frame (>5x) on SemanticKITTI and achieves superior performance compared to state-of-the-art 3D LiDAR scene completion models. Our model and code are publicly available on https://github.com/happyw1nd/ScoreLiDAR.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03515
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion
Zhang, Shengyuan
Zhao, An
Yang, Ling
Li, Zejian
Meng, Chenye
Xu, Haoran
Chen, Tianrun
Wei, AnYang
GU, Perry Pengyun
Sun, Lingyun
Computer Vision and Pattern Recognition
Diffusion models have been applied to 3D LiDAR scene completion due to their strong training stability and high completion quality. However, the slow sampling speed limits the practical application of diffusion-based scene completion models since autonomous vehicles require an efficient perception of surrounding environments. This paper proposes a novel distillation method tailored for 3D Li- DAR scene completion models, dubbed ScoreLiDAR, which achieves efficient yet high-quality scene completion. Score- LiDAR enables the distilled model to sample in significantly fewer steps after distillation. To improve completion quality, we also introduce a novel Structural Loss, which encourages the distilled model to capture the geometric structure of the 3D LiDAR scene. The loss contains a scene-wise term constraining the holistic structure and a point-wise term constraining the key landmark points and their relative configuration. Extensive experiments demonstrate that ScoreLiDAR significantly accelerates the completion time from 30.55 to 5.37 seconds per frame (>5x) on SemanticKITTI and achieves superior performance compared to state-of-the-art 3D LiDAR scene completion models. Our model and code are publicly available on https://github.com/happyw1nd/ScoreLiDAR.
title Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.03515