MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gidaris, Spyros, Bursuc, Andrei, Simeoni, Oriane, Vobecky, Antonin, Komodakis, Nikos, Cord, Matthieu, Pérez, Patrick
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914871686201344
author Gidaris, Spyros
Bursuc, Andrei
Simeoni, Oriane
Vobecky, Antonin
Komodakis, Nikos
Cord, Matthieu
Pérez, Patrick
author_facet Gidaris, Spyros
Bursuc, Andrei
Simeoni, Oriane
Vobecky, Antonin
Komodakis, Nikos
Cord, Matthieu
Pérez, Patrick
contents Self-supervised learning can be used for mitigating the greedy needs of Vision Transformer networks for very large fully-annotated datasets. Different classes of self-supervised learning offer representations with either good contextual reasoning properties, e.g., using masked image modeling strategies, or invariance to image perturbations, e.g., with contrastive methods. In this work, we propose a single-stage and standalone method, MOCA, which unifies both desired properties using novel mask-and-predict objectives defined with high-level features (instead of pixel-level details). Moreover, we show how to effectively employ both learning paradigms in a synergistic and computation-efficient way. Doing so, we achieve new state-of-the-art results on low-shot settings and strong experimental results in various evaluation protocols with a training that is at least 3 times faster than prior methods. We provide the implementation code at https://github.com/valeoai/MOCA.
format Preprint
id arxiv_https___arxiv_org_abs_2307_09361
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
Gidaris, Spyros
Bursuc, Andrei
Simeoni, Oriane
Vobecky, Antonin
Komodakis, Nikos
Cord, Matthieu
Pérez, Patrick
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Self-supervised learning can be used for mitigating the greedy needs of Vision Transformer networks for very large fully-annotated datasets. Different classes of self-supervised learning offer representations with either good contextual reasoning properties, e.g., using masked image modeling strategies, or invariance to image perturbations, e.g., with contrastive methods. In this work, we propose a single-stage and standalone method, MOCA, which unifies both desired properties using novel mask-and-predict objectives defined with high-level features (instead of pixel-level details). Moreover, we show how to effectively employ both learning paradigms in a synergistic and computation-efficient way. Doing so, we achieve new state-of-the-art results on low-shot settings and strong experimental results in various evaluation protocols with a training that is at least 3 times faster than prior methods. We provide the implementation code at https://github.com/valeoai/MOCA.
title MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2307.09361