S-JEA: Stacked Joint Embedding Architectures for Self-Supervised Visual Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Manová, Alžběta, Durrant, Aiden, Leontidis, Georgios
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913571365978112
author Manová, Alžběta
Durrant, Aiden
Leontidis, Georgios
author_facet Manová, Alžběta
Durrant, Aiden
Leontidis, Georgios
contents The recent emergence of Self-Supervised Learning (SSL) as a fundamental paradigm for learning image representations has, and continues to, demonstrate high empirical success in a variety of tasks. However, most SSL approaches fail to learn embeddings that capture hierarchical semantic concepts that are separable and interpretable. In this work, we aim to learn highly separable semantic hierarchical representations by stacking Joint Embedding Architectures (JEA) where higher-level JEAs are input with representations of lower-level JEA. This results in a representation space that exhibits distinct sub-categories of semantic concepts (e.g., model and colour of vehicles) in higher-level JEAs. We empirically show that representations from stacked JEA perform on a similar level as traditional JEA with comparative parameter counts and visualise the representation spaces to validate the semantic hierarchies.
format Preprint
id arxiv_https___arxiv_org_abs_2305_11701
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle S-JEA: Stacked Joint Embedding Architectures for Self-Supervised Visual Representation Learning
Manová, Alžběta
Durrant, Aiden
Leontidis, Georgios
Computer Vision and Pattern Recognition
Machine Learning
The recent emergence of Self-Supervised Learning (SSL) as a fundamental paradigm for learning image representations has, and continues to, demonstrate high empirical success in a variety of tasks. However, most SSL approaches fail to learn embeddings that capture hierarchical semantic concepts that are separable and interpretable. In this work, we aim to learn highly separable semantic hierarchical representations by stacking Joint Embedding Architectures (JEA) where higher-level JEAs are input with representations of lower-level JEA. This results in a representation space that exhibits distinct sub-categories of semantic concepts (e.g., model and colour of vehicles) in higher-level JEAs. We empirically show that representations from stacked JEA perform on a similar level as traditional JEA with comparative parameter counts and visualise the representation spaces to validate the semantic hierarchies.
title S-JEA: Stacked Joint Embedding Architectures for Self-Supervised Visual Representation Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2305.11701