Saved in:
Bibliographic Details
Main Authors: Liu, Jiaming, Kong, Linghe, Wu, Yue, Gong, Maoguo, Li, Hao, Miao, Qiguang, Ma, Wenping, Qin, Can
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.17547
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917803456462848
author Liu, Jiaming
Kong, Linghe
Wu, Yue
Gong, Maoguo
Li, Hao
Miao, Qiguang
Ma, Wenping
Qin, Can
author_facet Liu, Jiaming
Kong, Linghe
Wu, Yue
Gong, Maoguo
Li, Hao
Miao, Qiguang
Ma, Wenping
Qin, Can
contents Existing 3D mask learning methods encounter performance bottlenecks under limited data, and our objective is to overcome this limitation. In this paper, we introduce a triple point masking scheme, named TPM, which serves as a scalable framework for pre-training of masked autoencoders to achieve multi-mask learning for 3D point clouds. Specifically, we augment the baselines with two additional mask choices (i.e., medium mask and low mask) as our core insight is that the recovery process of an object can manifest in diverse ways. Previous high-masking schemes focus on capturing the global representation but lack the fine-grained recovery capability, so that the generated pre-trained weights tend to play a limited role in the fine-tuning process. With the support of the proposed TPM, available methods can exhibit more flexible and accurate completion capabilities, enabling the potential autoencoder in the pre-training stage to consider multiple representations of a single 3D object. In addition, an SVM-guided weight selection module is proposed to fill the encoder parameters for downstream networks with the optimal weight during the fine-tuning stage, maximizing linear accuracy and facilitating the acquisition of intricate representations for new objects. Extensive experiments show that the four baselines equipped with the proposed TPM achieve comprehensive performance improvements on various downstream tasks. Our code and models are available at https://github.com/liujia99/TPM.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17547
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Triple Point Masking
Liu, Jiaming
Kong, Linghe
Wu, Yue
Gong, Maoguo
Li, Hao
Miao, Qiguang
Ma, Wenping
Qin, Can
Computer Vision and Pattern Recognition
Artificial Intelligence
Existing 3D mask learning methods encounter performance bottlenecks under limited data, and our objective is to overcome this limitation. In this paper, we introduce a triple point masking scheme, named TPM, which serves as a scalable framework for pre-training of masked autoencoders to achieve multi-mask learning for 3D point clouds. Specifically, we augment the baselines with two additional mask choices (i.e., medium mask and low mask) as our core insight is that the recovery process of an object can manifest in diverse ways. Previous high-masking schemes focus on capturing the global representation but lack the fine-grained recovery capability, so that the generated pre-trained weights tend to play a limited role in the fine-tuning process. With the support of the proposed TPM, available methods can exhibit more flexible and accurate completion capabilities, enabling the potential autoencoder in the pre-training stage to consider multiple representations of a single 3D object. In addition, an SVM-guided weight selection module is proposed to fill the encoder parameters for downstream networks with the optimal weight during the fine-tuning stage, maximizing linear accuracy and facilitating the acquisition of intricate representations for new objects. Extensive experiments show that the four baselines equipped with the proposed TPM achieve comprehensive performance improvements on various downstream tasks. Our code and models are available at https://github.com/liujia99/TPM.
title Triple Point Masking
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2409.17547