On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation Paradigm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Peng, Shi, Bei, Yu, Daiwei, Lin, Tao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909141047443456
author Sun, Peng
Shi, Bei
Yu, Daiwei
Lin, Tao
author_facet Sun, Peng
Shi, Bei
Yu, Daiwei
Lin, Tao
contents Contemporary machine learning requires training large neural networks on massive datasets and thus faces the challenges of high computational demands. Dataset distillation, as a recent emerging strategy, aims to compress real-world datasets for efficient training. However, this line of research currently struggle with large-scale and high-resolution datasets, hindering its practicality and feasibility. To this end, we re-examine the existing dataset distillation methods and identify three properties required for large-scale real-world applications, namely, realism, diversity, and efficiency. As a remedy, we propose RDED, a novel computationally-efficient yet effective data distillation paradigm, to enable both diversity and realism of the distilled data. Extensive empirical results over various neural architectures and datasets demonstrate the advancement of RDED: we can distill the full ImageNet-1K to a small dataset comprising 10 images per class within 7 minutes, achieving a notable 42% top-1 accuracy with ResNet-18 on a single RTX-4090 GPU (while the SOTA only achieves 21% but requires 6 hours).
format Preprint
id arxiv_https___arxiv_org_abs_2312_03526
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation Paradigm
Sun, Peng
Shi, Bei
Yu, Daiwei
Lin, Tao
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Contemporary machine learning requires training large neural networks on massive datasets and thus faces the challenges of high computational demands. Dataset distillation, as a recent emerging strategy, aims to compress real-world datasets for efficient training. However, this line of research currently struggle with large-scale and high-resolution datasets, hindering its practicality and feasibility. To this end, we re-examine the existing dataset distillation methods and identify three properties required for large-scale real-world applications, namely, realism, diversity, and efficiency. As a remedy, we propose RDED, a novel computationally-efficient yet effective data distillation paradigm, to enable both diversity and realism of the distilled data. Extensive empirical results over various neural architectures and datasets demonstrate the advancement of RDED: we can distill the full ImageNet-1K to a small dataset comprising 10 images per class within 7 minutes, achieving a notable 42% top-1 accuracy with ResNet-18 on a single RTX-4090 GPU (while the SOTA only achieves 21% but requires 6 hours).
title On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation Paradigm
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.03526