Distill the Best, Ignore the Rest: Improving Dataset Distillation with Loss-Value-Based Pruning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Moser, Brian B., Raue, Federico, Nauen, Tobias C., Frolov, Stanislav, Dengel, Andreas
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909980435677184
author Moser, Brian B.
Raue, Federico
Nauen, Tobias C.
Frolov, Stanislav
Dengel, Andreas
author_facet Moser, Brian B.
Raue, Federico
Nauen, Tobias C.
Frolov, Stanislav
Dengel, Andreas
contents Dataset distillation has gained significant interest in recent years, yet existing approaches typically distill from the entire dataset, potentially including non-beneficial samples. We introduce a novel "Prune First, Distill After" framework that systematically prunes datasets via loss-based sampling prior to distillation. By leveraging pruning before classical distillation techniques and generative priors, we create a representative core-set that leads to enhanced generalization for unseen architectures - a significant challenge of current distillation methods. More specifically, our proposed framework significantly boosts distilled quality, achieving up to a 5.2 percentage points accuracy increase even with substantial dataset pruning, i.e., removing 80% of the original dataset prior to distillation. Overall, our experimental results highlight the advantages of our easy-sample prioritization and cross-architecture robustness, paving the way for more effective and high-quality dataset distillation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12115
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Distill the Best, Ignore the Rest: Improving Dataset Distillation with Loss-Value-Based Pruning
Moser, Brian B.
Raue, Federico
Nauen, Tobias C.
Frolov, Stanislav
Dengel, Andreas
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Dataset distillation has gained significant interest in recent years, yet existing approaches typically distill from the entire dataset, potentially including non-beneficial samples. We introduce a novel "Prune First, Distill After" framework that systematically prunes datasets via loss-based sampling prior to distillation. By leveraging pruning before classical distillation techniques and generative priors, we create a representative core-set that leads to enhanced generalization for unseen architectures - a significant challenge of current distillation methods. More specifically, our proposed framework significantly boosts distilled quality, achieving up to a 5.2 percentage points accuracy increase even with substantial dataset pruning, i.e., removing 80% of the original dataset prior to distillation. Overall, our experimental results highlight the advantages of our easy-sample prioritization and cross-architecture robustness, paving the way for more effective and high-quality dataset distillation.
title Distill the Best, Ignore the Rest: Improving Dataset Distillation with Loss-Value-Based Pruning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.12115