Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qu, Wentao, Mei, Guofeng, Wang, Jing, Wu, Yujiao, Huang, Xiaoshui, Xiao, Liang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915465162391552
author Qu, Wentao
Mei, Guofeng
Wang, Jing
Wu, Yujiao
Huang, Xiaoshui
Xiao, Liang
author_facet Qu, Wentao
Mei, Guofeng
Wang, Jing
Wu, Yujiao
Huang, Xiaoshui
Xiao, Liang
contents Denoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To address this, we propose a Robust single-stage fully Sparse 3D object Detection Network with a Detachable Latent Framework (DLF) of DDPMs, named RSDNet. Specifically, RSDNet learns the denoising process in latent feature spaces through lightweight denoising networks like multi-level denoising autoencoders (DAEs). This enables RSDNet to effectively understand scene distributions under multi-level perturbations, achieving robust and reliable detection. Meanwhile, we reformulate the noising and denoising mechanisms of DDPMs, enabling DLF to construct multi-type and multi-level noise samples and targets, enhancing RSDNet robustness to multiple perturbations. Furthermore, a semantic-geometric conditional guidance is introduced to perceive the object boundaries and shapes, alleviating the center feature missing problem in sparse representations, enabling RSDNet to perform in a fully sparse detection pipeline. Moreover, the detachable denoising network design of DLF enables RSDNet to perform single-step detection in inference, further enhancing detection efficiency. Extensive experiments on public benchmarks show that RSDNet can outperform existing methods, achieving state-of-the-art detection.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03252
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion
Qu, Wentao
Mei, Guofeng
Wang, Jing
Wu, Yujiao
Huang, Xiaoshui
Xiao, Liang
Computer Vision and Pattern Recognition
Denoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To address this, we propose a Robust single-stage fully Sparse 3D object Detection Network with a Detachable Latent Framework (DLF) of DDPMs, named RSDNet. Specifically, RSDNet learns the denoising process in latent feature spaces through lightweight denoising networks like multi-level denoising autoencoders (DAEs). This enables RSDNet to effectively understand scene distributions under multi-level perturbations, achieving robust and reliable detection. Meanwhile, we reformulate the noising and denoising mechanisms of DDPMs, enabling DLF to construct multi-type and multi-level noise samples and targets, enhancing RSDNet robustness to multiple perturbations. Furthermore, a semantic-geometric conditional guidance is introduced to perceive the object boundaries and shapes, alleviating the center feature missing problem in sparse representations, enabling RSDNet to perform in a fully sparse detection pipeline. Moreover, the detachable denoising network design of DLF enables RSDNet to perform single-step detection in inference, further enhancing detection efficiency. Extensive experiments on public benchmarks show that RSDNet can outperform existing methods, achieving state-of-the-art detection.
title Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.03252