Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vo, Huy V., Khalidov, Vasil, Darcet, Timothée, Moutakanni, Théo, Smetanin, Nikita, Szafraniec, Marc, Touvron, Hugo, Couprie, Camille, Oquab, Maxime, Joulin, Armand, Jégou, Hervé, Labatut, Patrick, Bojanowski, Piotr
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917707984666624
author Vo, Huy V.
Khalidov, Vasil
Darcet, Timothée
Moutakanni, Théo
Smetanin, Nikita
Szafraniec, Marc
Touvron, Hugo
Couprie, Camille
Oquab, Maxime
Joulin, Armand
Jégou, Hervé
Labatut, Patrick
Bojanowski, Piotr
author_facet Vo, Huy V.
Khalidov, Vasil
Darcet, Timothée
Moutakanni, Théo
Smetanin, Nikita
Szafraniec, Marc
Touvron, Hugo
Couprie, Camille
Oquab, Maxime
Joulin, Armand
Jégou, Hervé
Labatut, Patrick
Bojanowski, Piotr
contents Self-supervised features are the cornerstone of modern machine learning systems. They are typically pre-trained on data collections whose construction and curation typically require extensive human effort. This manual process has some limitations similar to those encountered in supervised learning, e.g., the crowd-sourced selection of data is costly and time-consuming, preventing scaling the dataset size. In this work, we consider the problem of automatic curation of high-quality datasets for self-supervised pre-training. We posit that such datasets should be large, diverse and balanced, and propose a clustering-based approach for building ones satisfying all these criteria. Our method involves successive and hierarchical applications of $k$-means on a large and diverse data repository to obtain clusters that distribute uniformly among data concepts, followed by a hierarchical, balanced sampling step from these clusters. Extensive experiments on three different data domains including web-based images, satellite images and text show that features trained on our automatically curated datasets outperform those trained on uncurated data while being on par or better than ones trained on manually curated data. Code is available at https://github.com/facebookresearch/ssl-data-curation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach
Vo, Huy V.
Khalidov, Vasil
Darcet, Timothée
Moutakanni, Théo
Smetanin, Nikita
Szafraniec, Marc
Touvron, Hugo
Couprie, Camille
Oquab, Maxime
Joulin, Armand
Jégou, Hervé
Labatut, Patrick
Bojanowski, Piotr
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Self-supervised features are the cornerstone of modern machine learning systems. They are typically pre-trained on data collections whose construction and curation typically require extensive human effort. This manual process has some limitations similar to those encountered in supervised learning, e.g., the crowd-sourced selection of data is costly and time-consuming, preventing scaling the dataset size. In this work, we consider the problem of automatic curation of high-quality datasets for self-supervised pre-training. We posit that such datasets should be large, diverse and balanced, and propose a clustering-based approach for building ones satisfying all these criteria. Our method involves successive and hierarchical applications of $k$-means on a large and diverse data repository to obtain clusters that distribute uniformly among data concepts, followed by a hierarchical, balanced sampling step from these clusters. Extensive experiments on three different data domains including web-based images, satellite images and text show that features trained on our automatically curated datasets outperform those trained on uncurated data while being on par or better than ones trained on manually curated data. Code is available at https://github.com/facebookresearch/ssl-data-curation.
title Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.15613