Beyond [cls]: Exploring the true potential of Masked Image Modeling representations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Przewięźlikowski, Marcin, Balestriero, Randall, Jasiński, Wojciech, Śmieja, Marek, Zieliński, Bartosz
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915550264819712
author Przewięźlikowski, Marcin
Balestriero, Randall
Jasiński, Wojciech
Śmieja, Marek
Zieliński, Bartosz
author_facet Przewięźlikowski, Marcin
Balestriero, Randall
Jasiński, Wojciech
Śmieja, Marek
Zieliński, Bartosz
contents Masked Image Modeling (MIM) has emerged as a promising approach for Self-Supervised Learning (SSL) of visual representations. However, the out-of-the-box performance of MIMs is typically inferior to competing approaches. Most users cannot afford fine-tuning due to the need for large amounts of data, high GPU consumption, and specialized user knowledge. Therefore, the practical use of MIM representations is limited. In this paper we ask what is the reason for the poor out-of-the-box performance of MIMs. Is it due to weaker features produced by MIM models, or is it due to suboptimal usage? Through detailed analysis, we show that attention in MIMs is spread almost uniformly over many patches, leading to ineffective aggregation by the [cls] token. Based on this insight, we propose Selective Aggregation to better capture the rich semantic information retained in patch tokens, which significantly improves the out-of-the-box performance of MIM.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03215
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond [cls]: Exploring the true potential of Masked Image Modeling representations
Przewięźlikowski, Marcin
Balestriero, Randall
Jasiński, Wojciech
Śmieja, Marek
Zieliński, Bartosz
Computer Vision and Pattern Recognition
Machine Learning
Masked Image Modeling (MIM) has emerged as a promising approach for Self-Supervised Learning (SSL) of visual representations. However, the out-of-the-box performance of MIMs is typically inferior to competing approaches. Most users cannot afford fine-tuning due to the need for large amounts of data, high GPU consumption, and specialized user knowledge. Therefore, the practical use of MIM representations is limited. In this paper we ask what is the reason for the poor out-of-the-box performance of MIMs. Is it due to weaker features produced by MIM models, or is it due to suboptimal usage? Through detailed analysis, we show that attention in MIMs is spread almost uniformly over many patches, leading to ineffective aggregation by the [cls] token. Based on this insight, we propose Selective Aggregation to better capture the rich semantic information retained in patch tokens, which significantly improves the out-of-the-box performance of MIM.
title Beyond [cls]: Exploring the true potential of Masked Image Modeling representations
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.03215