On the Discriminability of Self-Supervised Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Song, Zeen, Qiang, Wenwen, Zheng, Changwen, Sun, Fuchun, Xiong, Hui
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911089288019968
author Song, Zeen
Qiang, Wenwen
Zheng, Changwen
Sun, Fuchun
Xiong, Hui
author_facet Song, Zeen
Qiang, Wenwen
Zheng, Changwen
Sun, Fuchun
Xiong, Hui
contents Self-supervised learning (SSL) has recently shown notable success in various visual tasks. However, in terms of discriminability, SSL is still not on par with supervised learning (SL). This paper identifies a key issue, the ``crowding problem," where features from different classes are not well-separated, and there is high intra-class variance. In contrast, SL ensures clear class separation. Our analysis reveals that SSL objectives do not adequately constrain the relationships between samples and their augmentations, leading to poorer performance in complex tasks. We further establish a theoretical framework that connects SSL objectives to cross-entropy risk bounds, explaining how reducing intra-class variance and increasing inter-class separation can improve generalization. To address this, we propose the Dynamic Semantic Adjuster (DSA), a learnable regulator that enhances feature aggregation and separation while being robust to outliers. Comprehensive experiments conducted on diverse benchmark datasets validate that DSA leads to substantial gains in SSL performance, narrowing the performance gap with SL.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13541
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Discriminability of Self-Supervised Representation Learning
Song, Zeen
Qiang, Wenwen
Zheng, Changwen
Sun, Fuchun
Xiong, Hui
Computer Vision and Pattern Recognition
Self-supervised learning (SSL) has recently shown notable success in various visual tasks. However, in terms of discriminability, SSL is still not on par with supervised learning (SL). This paper identifies a key issue, the ``crowding problem," where features from different classes are not well-separated, and there is high intra-class variance. In contrast, SL ensures clear class separation. Our analysis reveals that SSL objectives do not adequately constrain the relationships between samples and their augmentations, leading to poorer performance in complex tasks. We further establish a theoretical framework that connects SSL objectives to cross-entropy risk bounds, explaining how reducing intra-class variance and increasing inter-class separation can improve generalization. To address this, we propose the Dynamic Semantic Adjuster (DSA), a learnable regulator that enhances feature aggregation and separation while being robust to outliers. Comprehensive experiments conducted on diverse benchmark datasets validate that DSA leads to substantial gains in SSL performance, narrowing the performance gap with SL.
title On the Discriminability of Self-Supervised Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.13541