Self-Guided Masked Autoencoder

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shin, Jeongwoo, Lee, Inseo, Lee, Junho, Lee, Joonseok
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918104838176768
author Shin, Jeongwoo
Lee, Inseo
Lee, Junho
Lee, Joonseok
author_facet Shin, Jeongwoo
Lee, Inseo
Lee, Junho
Lee, Joonseok
contents Masked Autoencoder (MAE) is a self-supervised approach for representation learning, widely applicable to a variety of downstream tasks in computer vision. In spite of its success, it is still not fully uncovered what and how MAE exactly learns. In this paper, with an in-depth analysis, we discover that MAE intrinsically learns pattern-based patch-level clustering from surprisingly early stages of pretraining. Upon this understanding, we propose self-guided masked autoencoder, which internally generates informed mask by utilizing its progress in patch clustering, substituting the naive random masking of the vanilla MAE. Our approach significantly boosts its learning process without relying on any external models or supplementary information, keeping the benefit of self-supervised nature of MAE intact. Comprehensive experiments on various downstream tasks verify the effectiveness of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Guided Masked Autoencoder
Shin, Jeongwoo
Lee, Inseo
Lee, Junho
Lee, Joonseok
Computer Vision and Pattern Recognition
Masked Autoencoder (MAE) is a self-supervised approach for representation learning, widely applicable to a variety of downstream tasks in computer vision. In spite of its success, it is still not fully uncovered what and how MAE exactly learns. In this paper, with an in-depth analysis, we discover that MAE intrinsically learns pattern-based patch-level clustering from surprisingly early stages of pretraining. Upon this understanding, we propose self-guided masked autoencoder, which internally generates informed mask by utilizing its progress in patch clustering, substituting the naive random masking of the vanilla MAE. Our approach significantly boosts its learning process without relying on any external models or supplementary information, keeping the benefit of self-supervised nature of MAE intact. Comprehensive experiments on various downstream tasks verify the effectiveness of the proposed method.
title Self-Guided Masked Autoencoder
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.19773