Dynamically Scaled Temperature in Self-Supervised Contrastive Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Manna, Siladittya, Chattopadhyay, Soumitri, Dey, Rakesh, Bhattacharya, Saumik, Pal, Umapada
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917662316036096
author Manna, Siladittya
Chattopadhyay, Soumitri
Dey, Rakesh
Bhattacharya, Saumik
Pal, Umapada
author_facet Manna, Siladittya
Chattopadhyay, Soumitri
Dey, Rakesh
Bhattacharya, Saumik
Pal, Umapada
contents In contemporary self-supervised contrastive algorithms like SimCLR, MoCo, etc., the task of balancing attraction between two semantically similar samples and repulsion between two samples of different classes is primarily affected by the presence of hard negative samples. While the InfoNCE loss has been shown to impose penalties based on hardness, the temperature hyper-parameter is the key to regulating the penalties and the trade-off between uniformity and tolerance. In this work, we focus our attention on improving the performance of InfoNCE loss in self-supervised learning by proposing a novel cosine similarity dependent temperature scaling function to effectively optimize the distribution of the samples in the feature space. We also provide mathematical analyses to support the construction of such a dynamically scaled temperature function. Experimental evidence shows that the proposed framework outperforms the contrastive loss-based SSL algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2308_01140
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dynamically Scaled Temperature in Self-Supervised Contrastive Learning
Manna, Siladittya
Chattopadhyay, Soumitri
Dey, Rakesh
Bhattacharya, Saumik
Pal, Umapada
Machine Learning
Computer Vision and Pattern Recognition
In contemporary self-supervised contrastive algorithms like SimCLR, MoCo, etc., the task of balancing attraction between two semantically similar samples and repulsion between two samples of different classes is primarily affected by the presence of hard negative samples. While the InfoNCE loss has been shown to impose penalties based on hardness, the temperature hyper-parameter is the key to regulating the penalties and the trade-off between uniformity and tolerance. In this work, we focus our attention on improving the performance of InfoNCE loss in self-supervised learning by proposing a novel cosine similarity dependent temperature scaling function to effectively optimize the distribution of the samples in the feature space. We also provide mathematical analyses to support the construction of such a dynamically scaled temperature function. Experimental evidence shows that the proposed framework outperforms the contrastive loss-based SSL algorithms.
title Dynamically Scaled Temperature in Self-Supervised Contrastive Learning
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.01140