SAS: Semantic-aware Sampling for Generative Dataset Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Mingzhuo, Li, Guang, Ye, Linfeng, Mao, Jiafeng, Ogawa, Takahiro, Plataniotis, Konstantinos N., Haseyama, Miki
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911693849755648
author Li, Mingzhuo
Li, Guang
Ye, Linfeng
Mao, Jiafeng
Ogawa, Takahiro
Plataniotis, Konstantinos N.
Haseyama, Miki
author_facet Li, Mingzhuo
Li, Guang
Ye, Linfeng
Mao, Jiafeng
Ogawa, Takahiro
Plataniotis, Konstantinos N.
Haseyama, Miki
contents Deep neural networks have achieved impressive performance across a wide range of tasks, but this success often comes with substantial computational and storage costs due to large-scale training data. Dataset distillation addresses this challenge by constructing compact yet informative datasets that enable efficient model training while maintaining downstream performance. However, most existing approaches primarily emphasize matching data distributions or downstream training statistics, with limited attention to preserving high-level semantic information in the distilled data. In this work, we introduce a semantic-aware perspective for dataset distillation by leveraging Contrastive Language-Image Pretraining (CLIP) as a semantic prior for post-sampling. Our goal is to obtain distilled datasets that are not only compact but also semantically class-discriminative and diverse. To this end, we design three semantic scoring functions that quantify class relevance, inter-class separability, and intra-set diversity in a pretrained semantic space. Based on image pools generated by existing distillation methods, we further develop a two-stage strategy for effective sampling: the first stage filters semantically discriminative samples to form a reliable candidate set, and the second stage performs a dynamic diversity-aware selection to reduce redundancy while preserving semantic coverage. Extensive experiments across multiple datasets, image pools, and downstream models demonstrate consistent performance gains, highlighting the effectiveness of incorporating semantic information into dataset distillation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18012
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SAS: Semantic-aware Sampling for Generative Dataset Distillation
Li, Mingzhuo
Li, Guang
Ye, Linfeng
Mao, Jiafeng
Ogawa, Takahiro
Plataniotis, Konstantinos N.
Haseyama, Miki
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Deep neural networks have achieved impressive performance across a wide range of tasks, but this success often comes with substantial computational and storage costs due to large-scale training data. Dataset distillation addresses this challenge by constructing compact yet informative datasets that enable efficient model training while maintaining downstream performance. However, most existing approaches primarily emphasize matching data distributions or downstream training statistics, with limited attention to preserving high-level semantic information in the distilled data. In this work, we introduce a semantic-aware perspective for dataset distillation by leveraging Contrastive Language-Image Pretraining (CLIP) as a semantic prior for post-sampling. Our goal is to obtain distilled datasets that are not only compact but also semantically class-discriminative and diverse. To this end, we design three semantic scoring functions that quantify class relevance, inter-class separability, and intra-set diversity in a pretrained semantic space. Based on image pools generated by existing distillation methods, we further develop a two-stage strategy for effective sampling: the first stage filters semantically discriminative samples to form a reliable candidate set, and the second stage performs a dynamic diversity-aware selection to reduce redundancy while preserving semantic coverage. Extensive experiments across multiple datasets, image pools, and downstream models demonstrate consistent performance gains, highlighting the effectiveness of incorporating semantic information into dataset distillation.
title SAS: Semantic-aware Sampling for Generative Dataset Distillation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.18012