Self-Supervised Dataset Distillation for Transfer Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Dong Bok, Lee, Seanie, Ko, Joonho, Kawaguchi, Kenji, Lee, Juho, Hwang, Sung Ju
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916201979969536
author Lee, Dong Bok
Lee, Seanie
Ko, Joonho
Kawaguchi, Kenji
Lee, Juho
Hwang, Sung Ju
author_facet Lee, Dong Bok
Lee, Seanie
Ko, Joonho
Kawaguchi, Kenji
Lee, Juho
Hwang, Sung Ju
contents Dataset distillation methods have achieved remarkable success in distilling a large dataset into a small set of representative samples. However, they are not designed to produce a distilled dataset that can be effectively used for facilitating self-supervised pre-training. To this end, we propose a novel problem of distilling an unlabeled dataset into a set of small synthetic samples for efficient self-supervised learning (SSL). We first prove that a gradient of synthetic samples with respect to a SSL objective in naive bilevel optimization is \textit{biased} due to the randomness originating from data augmentations or masking. To address this issue, we propose to minimize the mean squared error (MSE) between a model's representations of the synthetic examples and their corresponding learnable target feature representations for the inner objective, which does not introduce any randomness. Our primary motivation is that the model obtained by the proposed inner optimization can mimic the \textit{self-supervised target model}. To achieve this, we also introduce the MSE between representations of the inner model and the self-supervised target model on the original full dataset for outer optimization. Lastly, assuming that a feature extractor is fixed, we only optimize a linear head on top of the feature extractor, which allows us to reduce the computational cost and obtain a closed-form solution of the head with kernel ridge regression. We empirically validate the effectiveness of our method on various applications involving transfer learning.
format Preprint
id arxiv_https___arxiv_org_abs_2310_06511
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Self-Supervised Dataset Distillation for Transfer Learning
Lee, Dong Bok
Lee, Seanie
Ko, Joonho
Kawaguchi, Kenji
Lee, Juho
Hwang, Sung Ju
Machine Learning
Dataset distillation methods have achieved remarkable success in distilling a large dataset into a small set of representative samples. However, they are not designed to produce a distilled dataset that can be effectively used for facilitating self-supervised pre-training. To this end, we propose a novel problem of distilling an unlabeled dataset into a set of small synthetic samples for efficient self-supervised learning (SSL). We first prove that a gradient of synthetic samples with respect to a SSL objective in naive bilevel optimization is \textit{biased} due to the randomness originating from data augmentations or masking. To address this issue, we propose to minimize the mean squared error (MSE) between a model's representations of the synthetic examples and their corresponding learnable target feature representations for the inner objective, which does not introduce any randomness. Our primary motivation is that the model obtained by the proposed inner optimization can mimic the \textit{self-supervised target model}. To achieve this, we also introduce the MSE between representations of the inner model and the self-supervised target model on the original full dataset for outer optimization. Lastly, assuming that a feature extractor is fixed, we only optimize a linear head on top of the feature extractor, which allows us to reduce the computational cost and obtain a closed-form solution of the head with kernel ridge regression. We empirically validate the effectiveness of our method on various applications involving transfer learning.
title Self-Supervised Dataset Distillation for Transfer Learning
topic Machine Learning
url https://arxiv.org/abs/2310.06511