An Empirical Study of Self-supervised Learning with Wasserstein Distance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yamada, Makoto, Takezawa, Yuki, Houry, Guillaume, Dusterwald, Kira Michaela, Sulem, Deborah, Zhao, Han, Tsai, Yao-Hung Hubert
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929234595807232
author Yamada, Makoto
Takezawa, Yuki
Houry, Guillaume
Dusterwald, Kira Michaela
Sulem, Deborah
Zhao, Han
Tsai, Yao-Hung Hubert
author_facet Yamada, Makoto
Takezawa, Yuki
Houry, Guillaume
Dusterwald, Kira Michaela
Sulem, Deborah
Zhao, Han
Tsai, Yao-Hung Hubert
contents In this study, we delve into the problem of self-supervised learning (SSL) utilizing the 1-Wasserstein distance on a tree structure (a.k.a., Tree-Wasserstein distance (TWD)), where TWD is defined as the L1 distance between two tree-embedded vectors. In SSL methods, the cosine similarity is often utilized as an objective function; however, it has not been well studied when utilizing the Wasserstein distance. Training the Wasserstein distance is numerically challenging. Thus, this study empirically investigates a strategy for optimizing the SSL with the Wasserstein distance and finds a stable training procedure. More specifically, we evaluate the combination of two types of TWD (total variation and ClusterTree) and several probability models, including the softmax function, the ArcFace probability model, and simplicial embedding. We propose a simple yet effective Jeffrey divergence-based regularization method to stabilize optimization. Through empirical experiments on STL10, CIFAR10, CIFAR100, and SVHN, we find that a simple combination of the softmax function and TWD can obtain significantly lower results than the standard SimCLR. Moreover, a simple combination of TWD and SimSiam fails to train the model. We find that the model performance depends on the combination of TWD and probability model, and that the Jeffrey divergence regularization helps in model training. Finally, we show that the appropriate combination of the TWD and probability model outperforms cosine similarity-based representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2310_10143
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle An Empirical Study of Self-supervised Learning with Wasserstein Distance
Yamada, Makoto
Takezawa, Yuki
Houry, Guillaume
Dusterwald, Kira Michaela
Sulem, Deborah
Zhao, Han
Tsai, Yao-Hung Hubert
Machine Learning
In this study, we delve into the problem of self-supervised learning (SSL) utilizing the 1-Wasserstein distance on a tree structure (a.k.a., Tree-Wasserstein distance (TWD)), where TWD is defined as the L1 distance between two tree-embedded vectors. In SSL methods, the cosine similarity is often utilized as an objective function; however, it has not been well studied when utilizing the Wasserstein distance. Training the Wasserstein distance is numerically challenging. Thus, this study empirically investigates a strategy for optimizing the SSL with the Wasserstein distance and finds a stable training procedure. More specifically, we evaluate the combination of two types of TWD (total variation and ClusterTree) and several probability models, including the softmax function, the ArcFace probability model, and simplicial embedding. We propose a simple yet effective Jeffrey divergence-based regularization method to stabilize optimization. Through empirical experiments on STL10, CIFAR10, CIFAR100, and SVHN, we find that a simple combination of the softmax function and TWD can obtain significantly lower results than the standard SimCLR. Moreover, a simple combination of TWD and SimSiam fails to train the model. We find that the model performance depends on the combination of TWD and probability model, and that the Jeffrey divergence regularization helps in model training. Finally, we show that the appropriate combination of the TWD and probability model outperforms cosine similarity-based representation learning.
title An Empirical Study of Self-supervised Learning with Wasserstein Distance
topic Machine Learning
url https://arxiv.org/abs/2310.10143