Towards Principled Dataset Distillation: A Spectral Distribution Perspective

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Ruixi, Wang, Shaobo, Chen, Jiahuan, Liu, Zhiyuan, Yang, Yicun, Chen, Zhaorun, Li, Zekai, Li, Kaixin, Wang, Xinming, Yi, Hongzhu, Wang, Kai, Zhang, Linfeng
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908860239839232
author Wu, Ruixi
Wang, Shaobo
Chen, Jiahuan
Liu, Zhiyuan
Yang, Yicun
Chen, Zhaorun
Li, Zekai
Li, Kaixin
Wang, Xinming
Yi, Hongzhu
Wang, Kai
Zhang, Linfeng
author_facet Wu, Ruixi
Wang, Shaobo
Chen, Jiahuan
Liu, Zhiyuan
Yang, Yicun
Chen, Zhaorun
Li, Zekai
Li, Kaixin
Wang, Xinming
Yi, Hongzhu
Wang, Kai
Zhang, Linfeng
contents Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic counterparts for efficient model training. However, existing DD methods exhibit substantial performance degradation on long-tailed datasets. We identify two fundamental challenges: heuristic design choices for distribution discrepancy measure and uniform treatment of imbalanced classes. To address these limitations, we propose Class-Aware Spectral Distribution Matching (CSDM), which reformulates distribution alignment via the spectrum of a well-behaved kernel function. This technique maps the original samples into frequency space, resulting in the Spectral Distribution Distance (SDD). To mitigate class imbalance, we exploit the unified form of SDD to perform amplitude-phase decomposition, which adaptively prioritizes the realism in tail classes. On CIFAR-10-LT, with 10 images per class, CSDM achieves a 14.0% improvement over state-of-the-art DD methods, with only a 5.7% performance drop when the number of images in tail classes decreases from 500 to 25, demonstrating strong stability on long-tailed data.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01698
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Principled Dataset Distillation: A Spectral Distribution Perspective
Wu, Ruixi
Wang, Shaobo
Chen, Jiahuan
Liu, Zhiyuan
Yang, Yicun
Chen, Zhaorun
Li, Zekai
Li, Kaixin
Wang, Xinming
Yi, Hongzhu
Wang, Kai
Zhang, Linfeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic counterparts for efficient model training. However, existing DD methods exhibit substantial performance degradation on long-tailed datasets. We identify two fundamental challenges: heuristic design choices for distribution discrepancy measure and uniform treatment of imbalanced classes. To address these limitations, we propose Class-Aware Spectral Distribution Matching (CSDM), which reformulates distribution alignment via the spectrum of a well-behaved kernel function. This technique maps the original samples into frequency space, resulting in the Spectral Distribution Distance (SDD). To mitigate class imbalance, we exploit the unified form of SDD to perform amplitude-phase decomposition, which adaptively prioritizes the realism in tail classes. On CIFAR-10-LT, with 10 images per class, CSDM achieves a 14.0% improvement over state-of-the-art DD methods, with only a 5.7% performance drop when the number of images in tail classes decreases from 500 to 25, demonstrating strong stability on long-tailed data.
title Towards Principled Dataset Distillation: A Spectral Distribution Perspective
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.01698