T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wei, Weijie, Nejadasl, Fatemeh Karimi, Gevers, Theo, Oswald, Martin R.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909262612004864
author Wei, Weijie
Nejadasl, Fatemeh Karimi
Gevers, Theo
Oswald, Martin R.
author_facet Wei, Weijie
Nejadasl, Fatemeh Karimi
Gevers, Theo
Oswald, Martin R.
contents The scarcity of annotated data in LiDAR point cloud understanding hinders effective representation learning. Consequently, scholars have been actively investigating efficacious self-supervised pre-training paradigms. Nevertheless, temporal information, which is inherent in the LiDAR point cloud sequence, is consistently disregarded. To better utilize this property, we propose an effective pre-training strategy, namely Temporal Masked Auto-Encoders (T-MAE), which takes as input temporally adjacent frames and learns temporal dependency. A SiamWCA backbone, containing a Siamese encoder and a windowed cross-attention (WCA) module, is established for the two-frame input. Considering that the movement of an ego-vehicle alters the view of the same instance, temporal modeling also serves as a robust and natural data augmentation, enhancing the comprehension of target objects. SiamWCA is a powerful architecture but heavily relies on annotated data. Our T-MAE pre-training strategy alleviates its demand for annotated data. Comprehensive experiments demonstrate that T-MAE achieves the best performance on both Waymo and ONCE datasets among competitive self-supervised approaches. Codes will be released at https://github.com/codename1995/T-MAE
format Preprint
id arxiv_https___arxiv_org_abs_2312_10217
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
Wei, Weijie
Nejadasl, Fatemeh Karimi
Gevers, Theo
Oswald, Martin R.
Computer Vision and Pattern Recognition
The scarcity of annotated data in LiDAR point cloud understanding hinders effective representation learning. Consequently, scholars have been actively investigating efficacious self-supervised pre-training paradigms. Nevertheless, temporal information, which is inherent in the LiDAR point cloud sequence, is consistently disregarded. To better utilize this property, we propose an effective pre-training strategy, namely Temporal Masked Auto-Encoders (T-MAE), which takes as input temporally adjacent frames and learns temporal dependency. A SiamWCA backbone, containing a Siamese encoder and a windowed cross-attention (WCA) module, is established for the two-frame input. Considering that the movement of an ego-vehicle alters the view of the same instance, temporal modeling also serves as a robust and natural data augmentation, enhancing the comprehension of target objects. SiamWCA is a powerful architecture but heavily relies on annotated data. Our T-MAE pre-training strategy alleviates its demand for annotated data. Comprehensive experiments demonstrate that T-MAE achieves the best performance on both Waymo and ONCE datasets among competitive self-supervised approaches. Codes will be released at https://github.com/codename1995/T-MAE
title T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.10217