TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Xiaopei, Hou, Yuenan, Huang, Xiaoshui, Lin, Binbin, He, Tong, Zhu, Xinge, Ma, Yuexin, Wu, Boxi, Liu, Haifeng, Cai, Deng, Ouyang, Wanli
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913431320264704
author Wu, Xiaopei
Hou, Yuenan
Huang, Xiaoshui
Lin, Binbin
He, Tong
Zhu, Xinge
Ma, Yuexin
Wu, Boxi
Liu, Haifeng
Cai, Deng
Ouyang, Wanli
author_facet Wu, Xiaopei
Hou, Yuenan
Huang, Xiaoshui
Lin, Binbin
He, Tong
Zhu, Xinge
Ma, Yuexin
Wu, Boxi
Liu, Haifeng
Cai, Deng
Ouyang, Wanli
contents Training deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the sparsity problem as it makes the input signal denser. However, previous multi-frame fusion algorithms fall short in utilizing sufficient temporal information due to the memory constraint, and they also ignore the informative temporal images. To fully exploit rich information hidden in long-term temporal point clouds and images, we present the Temporal Aggregation Network, termed TASeg. Specifically, we propose a Temporal LiDAR Aggregation and Distillation (TLAD) algorithm, which leverages historical priors to assign different aggregation steps for different classes. It can largely reduce memory and time overhead while achieving higher accuracy. Besides, TLAD trains a teacher injected with gt priors to distill the model, further boosting the performance. To make full use of temporal images, we design a Temporal Image Aggregation and Fusion (TIAF) module, which can greatly expand the camera FOV and enhance the present features. Temporal LiDAR points in the camera FOV are used as mediums to transform temporal image features to the present coordinate for temporal multi-modal fusion. Moreover, we develop a Static-Moving Switch Augmentation (SMSA) algorithm, which utilizes sufficient temporal information to enable objects to switch their motion states freely, thus greatly increasing static and moving training samples. Our TASeg ranks 1st on three challenging tracks, i.e., SemanticKITTI single-scan track, multi-scan track and nuScenes LiDAR segmentation track, strongly demonstrating the superiority of our method. Codes are available at https://github.com/LittlePey/TASeg.
format Preprint
id arxiv_https___arxiv_org_abs_2407_09751
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation
Wu, Xiaopei
Hou, Yuenan
Huang, Xiaoshui
Lin, Binbin
He, Tong
Zhu, Xinge
Ma, Yuexin
Wu, Boxi
Liu, Haifeng
Cai, Deng
Ouyang, Wanli
Computer Vision and Pattern Recognition
Training deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the sparsity problem as it makes the input signal denser. However, previous multi-frame fusion algorithms fall short in utilizing sufficient temporal information due to the memory constraint, and they also ignore the informative temporal images. To fully exploit rich information hidden in long-term temporal point clouds and images, we present the Temporal Aggregation Network, termed TASeg. Specifically, we propose a Temporal LiDAR Aggregation and Distillation (TLAD) algorithm, which leverages historical priors to assign different aggregation steps for different classes. It can largely reduce memory and time overhead while achieving higher accuracy. Besides, TLAD trains a teacher injected with gt priors to distill the model, further boosting the performance. To make full use of temporal images, we design a Temporal Image Aggregation and Fusion (TIAF) module, which can greatly expand the camera FOV and enhance the present features. Temporal LiDAR points in the camera FOV are used as mediums to transform temporal image features to the present coordinate for temporal multi-modal fusion. Moreover, we develop a Static-Moving Switch Augmentation (SMSA) algorithm, which utilizes sufficient temporal information to enable objects to switch their motion states freely, thus greatly increasing static and moving training samples. Our TASeg ranks 1st on three challenging tracks, i.e., SemanticKITTI single-scan track, multi-scan track and nuScenes LiDAR segmentation track, strongly demonstrating the superiority of our method. Codes are available at https://github.com/LittlePey/TASeg.
title TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.09751