On Partial Prototype Collapse in the DINO Family of Self-Supervised Methods

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Govindarajan, Hariprasath, Sidén, Per, Roll, Jacob, Lindsten, Fredrik
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917808096411648
author Govindarajan, Hariprasath
Sidén, Per
Roll, Jacob
Lindsten, Fredrik
author_facet Govindarajan, Hariprasath
Sidén, Per
Roll, Jacob
Lindsten, Fredrik
contents A prominent self-supervised learning paradigm is to model the representations as clusters, or more generally as a mixture model. Learning to map the data samples to compact representations and fitting the mixture model simultaneously leads to the representation collapse problem. Regularizing the distribution of data points over the clusters is the prevalent strategy to avoid this issue. While this is sufficient to prevent full representation collapse, we show that a partial prototype collapse problem still exists in the DINO family of methods, that leads to significant redundancies in the prototypes. Such prototype redundancies serve as shortcuts for the method to achieve a marginal latent class distribution that matches the prescribed prior. We show that by encouraging the model to use diverse prototypes, the partial prototype collapse can be mitigated. Effective utilization of the prototypes enables the methods to learn more fine-grained clusters, encouraging more informative representations. We demonstrate that this is especially beneficial when pre-training on a long-tailed fine-grained dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14060
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Partial Prototype Collapse in the DINO Family of Self-Supervised Methods
Govindarajan, Hariprasath
Sidén, Per
Roll, Jacob
Lindsten, Fredrik
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
A prominent self-supervised learning paradigm is to model the representations as clusters, or more generally as a mixture model. Learning to map the data samples to compact representations and fitting the mixture model simultaneously leads to the representation collapse problem. Regularizing the distribution of data points over the clusters is the prevalent strategy to avoid this issue. While this is sufficient to prevent full representation collapse, we show that a partial prototype collapse problem still exists in the DINO family of methods, that leads to significant redundancies in the prototypes. Such prototype redundancies serve as shortcuts for the method to achieve a marginal latent class distribution that matches the prescribed prior. We show that by encouraging the model to use diverse prototypes, the partial prototype collapse can be mitigated. Effective utilization of the prototypes enables the methods to learn more fine-grained clusters, encouraging more informative representations. We demonstrate that this is especially beneficial when pre-training on a long-tailed fine-grained dataset.
title On Partial Prototype Collapse in the DINO Family of Self-Supervised Methods
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.14060