IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Chenru, Chen, Yunyi, Yang, Zijun, Zhou, Joey Tianyi, Zhang, Chi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908886264446976
author Wang, Chenru
Chen, Yunyi
Yang, Zijun
Zhou, Joey Tianyi
Zhang, Chi
author_facet Wang, Chenru
Chen, Yunyi
Yang, Zijun
Zhou, Joey Tianyi
Zhang, Chi
contents Dataset Distillation aims to synthesize compact datasets that can approximate the training efficacy of large-scale real datasets, offering an efficient solution to the increasing computational demands of modern deep learning. Recently, diffusion-based dataset distillation methods have shown great promise by leveraging the strong generative capacity of diffusion models to produce diverse and structurally consistent samples. However, a fundamental goal misalignment persists: diffusion models are optimized for generative likelihood rather than discriminative utility, resulting in over-concentration in high-density regions and inadequate coverage of boundary samples crucial for classification. To address this issue, we propose two complementary strategies. Inversion-Matching (IM) introduces an inversion-guided fine-tuning process that aligns denoising trajectories with their inversion counterparts, broadening distributional coverage and enhancing diversity. Selective Subgroup Sampling(S^3) is a training-free sampling mechanism that improves inter-class separability by selecting synthetic subsets that are both representative and distinctive. Extensive experiments demonstrate that our approach significantly enhances the discriminative quality and generalization of distilled datasets, achieving state-of-the-art performance among diffusion-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13960
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation
Wang, Chenru
Chen, Yunyi
Yang, Zijun
Zhou, Joey Tianyi
Zhang, Chi
Computer Vision and Pattern Recognition
Dataset Distillation aims to synthesize compact datasets that can approximate the training efficacy of large-scale real datasets, offering an efficient solution to the increasing computational demands of modern deep learning. Recently, diffusion-based dataset distillation methods have shown great promise by leveraging the strong generative capacity of diffusion models to produce diverse and structurally consistent samples. However, a fundamental goal misalignment persists: diffusion models are optimized for generative likelihood rather than discriminative utility, resulting in over-concentration in high-density regions and inadequate coverage of boundary samples crucial for classification. To address this issue, we propose two complementary strategies. Inversion-Matching (IM) introduces an inversion-guided fine-tuning process that aligns denoising trajectories with their inversion counterparts, broadening distributional coverage and enhancing diversity. Selective Subgroup Sampling(S^3) is a training-free sampling mechanism that improves inter-class separability by selecting synthetic subsets that are both representative and distinctive. Extensive experiments demonstrate that our approach significantly enhances the discriminative quality and generalization of distilled datasets, achieving state-of-the-art performance among diffusion-based methods.
title IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.13960