Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zha, Yixin, Wang, Chuxin, Yang, Wenfei, Zhang, Tianzhu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911025989681152
author Zha, Yixin
Wang, Chuxin
Yang, Wenfei
Zhang, Tianzhu
author_facet Zha, Yixin
Wang, Chuxin
Yang, Wenfei
Zhang, Tianzhu
contents Point cloud understanding aims to acquire robust and general feature representations from unlabeled data. Masked point modeling-based methods have recently shown significant performance across various downstream tasks. These pre-training methods rely on random masking strategies to establish the perception of point clouds by restoring corrupted point cloud inputs, which leads to the failure of capturing reasonable semantic relationships by the self-supervised models. To address this issue, we propose Semantic Masked Autoencoder, which comprises two main components: a prototype-based component semantic modeling module and a component semantic-enhanced masking strategy. Specifically, in the component semantic modeling module, we design a component semantic guidance mechanism to direct a set of learnable prototypes in capturing the semantics of different components from objects. Leveraging these prototypes, we develop a component semantic-enhanced masking strategy that addresses the limitations of random masking in effectively covering complete component structures. Furthermore, we introduce a component semantic-enhanced prompt-tuning strategy, which further leverages these prototypes to improve the performance of pre-trained models in downstream tasks. Extensive experiments conducted on datasets such as ScanObjectNN, ModelNet40, and ShapeNetPart demonstrate the effectiveness of our proposed modules.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21957
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding
Zha, Yixin
Wang, Chuxin
Yang, Wenfei
Zhang, Tianzhu
Computer Vision and Pattern Recognition
Point cloud understanding aims to acquire robust and general feature representations from unlabeled data. Masked point modeling-based methods have recently shown significant performance across various downstream tasks. These pre-training methods rely on random masking strategies to establish the perception of point clouds by restoring corrupted point cloud inputs, which leads to the failure of capturing reasonable semantic relationships by the self-supervised models. To address this issue, we propose Semantic Masked Autoencoder, which comprises two main components: a prototype-based component semantic modeling module and a component semantic-enhanced masking strategy. Specifically, in the component semantic modeling module, we design a component semantic guidance mechanism to direct a set of learnable prototypes in capturing the semantics of different components from objects. Leveraging these prototypes, we develop a component semantic-enhanced masking strategy that addresses the limitations of random masking in effectively covering complete component structures. Furthermore, we introduce a component semantic-enhanced prompt-tuning strategy, which further leverages these prototypes to improve the performance of pre-trained models in downstream tasks. Extensive experiments conducted on datasets such as ScanObjectNN, ModelNet40, and ShapeNetPart demonstrate the effectiveness of our proposed modules.
title Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21957