The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Mo, Shentong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913623359619072
author Mo, Shentong
author_facet Mo, Shentong
contents Masked autoencoders (MAE) have recently succeeded in self-supervised vision representation learning. Previous work mainly applied custom-designed (e.g., random, block-wise) masking or teacher (e.g., CLIP)-guided masking and targets. However, they ignore the potential role of the self-training (student) model in giving feedback to the teacher for masking and targets. In this work, we present to integrate Collaborative Masking and Targets for boosting Masked AutoEncoders, namely CMT-MAE. Specifically, CMT-MAE leverages a simple collaborative masking mechanism through linear aggregation across attentions from both teacher and student models. We further propose using the output features from those two models as the collaborative target of the decoder. Our simple and effective framework pre-trained on ImageNet-1K achieves state-of-the-art linear probing and fine-tuning performance. In particular, using ViT-base, we improve the fine-tuning results of the vanilla MAE from 83.6% to 85.7%.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17566
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
Mo, Shentong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image and Video Processing
Signal Processing
Masked autoencoders (MAE) have recently succeeded in self-supervised vision representation learning. Previous work mainly applied custom-designed (e.g., random, block-wise) masking or teacher (e.g., CLIP)-guided masking and targets. However, they ignore the potential role of the self-training (student) model in giving feedback to the teacher for masking and targets. In this work, we present to integrate Collaborative Masking and Targets for boosting Masked AutoEncoders, namely CMT-MAE. Specifically, CMT-MAE leverages a simple collaborative masking mechanism through linear aggregation across attentions from both teacher and student models. We further propose using the output features from those two models as the collaborative target of the decoder. Our simple and effective framework pre-trained on ImageNet-1K achieves state-of-the-art linear probing and fine-tuning performance. In particular, using ViT-base, we improve the fine-tuning results of the vanilla MAE from 83.6% to 85.7%.
title The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image and Video Processing
Signal Processing
url https://arxiv.org/abs/2412.17566