SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yan, Xiangchao, Chen, Runjian, Zhang, Bo, Ye, Hancheng, Xia, Renqiu, Yuan, Jiakang, Zhou, Hongbin, Cai, Xinyu, Shi, Botian, Shao, Wenqi, Luo, Ping, Qiao, Yu, Chen, Tao, Yan, Junchi
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911059611222016
author Yan, Xiangchao
Chen, Runjian
Zhang, Bo
Ye, Hancheng
Xia, Renqiu
Yuan, Jiakang
Zhou, Hongbin
Cai, Xinyu
Shi, Botian
Shao, Wenqi
Luo, Ping
Qiao, Yu
Chen, Tao
Yan, Junchi
author_facet Yan, Xiangchao
Chen, Runjian
Zhang, Bo
Ye, Hancheng
Xia, Renqiu
Yuan, Jiakang
Zhou, Hongbin
Cai, Xinyu
Shi, Botian
Shao, Wenqi
Luo, Ping
Qiao, Yu
Chen, Tao
Yan, Junchi
contents Annotating 3D LiDAR point clouds for perception tasks is fundamental for many applications e.g., autonomous driving, yet it still remains notoriously labor-intensive. Pretraining-finetuning approach can alleviate the labeling burden by fine-tuning a pre-trained backbone across various downstream datasets as well as tasks. In this paper, we propose SPOT, namely Scalable Pre-training via Occupancy prediction for learning Transferable 3D representations under such a label-efficient fine-tuning paradigm. SPOT achieves effectiveness on various public datasets with different downstream tasks, showcasing its general representation power, cross-domain robustness and data scalability which are three key factors for real-world application. Specifically, we both theoretically and empirically show, for the first time, that general representations learning can be achieved through the task of occupancy prediction. Then, to address the domain gap caused by different LiDAR sensors and annotation methods, we develop a beam re-sampling technique for point cloud augmentation combined with class-balancing strategy. Furthermore, scalable pre-training is observed, that is, the downstream performance across all the experiments gets better with more pre-training data. Additionally, such pre-training strategy also remains compatible with unlabeled data. The hope is that our findings will facilitate the understanding of LiDAR points and pave the way for future advancements in LiDAR pre-training.
format Preprint
id arxiv_https___arxiv_org_abs_2309_10527
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
Yan, Xiangchao
Chen, Runjian
Zhang, Bo
Ye, Hancheng
Xia, Renqiu
Yuan, Jiakang
Zhou, Hongbin
Cai, Xinyu
Shi, Botian
Shao, Wenqi
Luo, Ping
Qiao, Yu
Chen, Tao
Yan, Junchi
Computer Vision and Pattern Recognition
Annotating 3D LiDAR point clouds for perception tasks is fundamental for many applications e.g., autonomous driving, yet it still remains notoriously labor-intensive. Pretraining-finetuning approach can alleviate the labeling burden by fine-tuning a pre-trained backbone across various downstream datasets as well as tasks. In this paper, we propose SPOT, namely Scalable Pre-training via Occupancy prediction for learning Transferable 3D representations under such a label-efficient fine-tuning paradigm. SPOT achieves effectiveness on various public datasets with different downstream tasks, showcasing its general representation power, cross-domain robustness and data scalability which are three key factors for real-world application. Specifically, we both theoretically and empirically show, for the first time, that general representations learning can be achieved through the task of occupancy prediction. Then, to address the domain gap caused by different LiDAR sensors and annotation methods, we develop a beam re-sampling technique for point cloud augmentation combined with class-balancing strategy. Furthermore, scalable pre-training is observed, that is, the downstream performance across all the experiments gets better with more pre-training data. Additionally, such pre-training strategy also remains compatible with unlabeled data. The hope is that our findings will facilitate the understanding of LiDAR points and pave the way for future advancements in LiDAR pre-training.
title SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.10527