Data-efficient Event Camera Pre-training via Disentangled Masked Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zhenpeng, Li, Chao, Chen, Hao, Deng, Yongjian, Geng, Yifeng, Wang, Limin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913250196586496
author Huang, Zhenpeng
Li, Chao
Chen, Hao
Deng, Yongjian
Geng, Yifeng
Wang, Limin
author_facet Huang, Zhenpeng
Li, Chao
Chen, Hao
Deng, Yongjian
Geng, Yifeng
Wang, Limin
contents In this paper, we present a new data-efficient voxel-based self-supervised learning method for event cameras. Our pre-training overcomes the limitations of previous methods, which either sacrifice temporal information by converting event sequences into 2D images for utilizing pre-trained image models or directly employ paired image data for knowledge distillation to enhance the learning of event streams. In order to make our pre-training data-efficient, we first design a semantic-uniform masking method to address the learning imbalance caused by the varying reconstruction difficulties of different regions in non-uniform data when using random masking. In addition, we ease the traditional hybrid masked modeling process by explicitly decomposing it into two branches, namely local spatio-temporal reconstruction and global semantic reconstruction to encourage the encoder to capture local correlations and global semantics, respectively. This decomposition allows our selfsupervised learning method to converge faster with minimal pre-training data. Compared to previous approaches, our self-supervised learning method does not rely on paired RGB images, yet enables simultaneous exploration of spatial and temporal cues in multiple scales. It exhibits excellent generalization performance and demonstrates significant improvements across various tasks with fewer parameters and lower computational costs.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00416
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data-efficient Event Camera Pre-training via Disentangled Masked Modeling
Huang, Zhenpeng
Li, Chao
Chen, Hao
Deng, Yongjian
Geng, Yifeng
Wang, Limin
Computer Vision and Pattern Recognition
In this paper, we present a new data-efficient voxel-based self-supervised learning method for event cameras. Our pre-training overcomes the limitations of previous methods, which either sacrifice temporal information by converting event sequences into 2D images for utilizing pre-trained image models or directly employ paired image data for knowledge distillation to enhance the learning of event streams. In order to make our pre-training data-efficient, we first design a semantic-uniform masking method to address the learning imbalance caused by the varying reconstruction difficulties of different regions in non-uniform data when using random masking. In addition, we ease the traditional hybrid masked modeling process by explicitly decomposing it into two branches, namely local spatio-temporal reconstruction and global semantic reconstruction to encourage the encoder to capture local correlations and global semantics, respectively. This decomposition allows our selfsupervised learning method to converge faster with minimal pre-training data. Compared to previous approaches, our self-supervised learning method does not rely on paired RGB images, yet enables simultaneous exploration of spatial and temporal cues in multiple scales. It exhibits excellent generalization performance and demonstrates significant improvements across various tasks with fewer parameters and lower computational costs.
title Data-efficient Event Camera Pre-training via Disentangled Masked Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.00416