MGD$^3$: Mode-Guided Dataset Distillation using Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chan-Santiago, Jeffrey A., Tirupattur, Praveen, Nayak, Gaurav Kumar, Liu, Gaowen, Shah, Mubarak
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910967427760128
author Chan-Santiago, Jeffrey A.
Tirupattur, Praveen
Nayak, Gaurav Kumar
Liu, Gaowen
Shah, Mubarak
author_facet Chan-Santiago, Jeffrey A.
Tirupattur, Praveen
Nayak, Gaurav Kumar
Liu, Gaowen
Shah, Mubarak
contents Dataset distillation has emerged as an effective strategy, significantly reducing training costs and facilitating more efficient model deployment. Recent advances have leveraged generative models to distill datasets by capturing the underlying data distribution. Unfortunately, existing methods require model fine-tuning with distillation losses to encourage diversity and representativeness. However, these methods do not guarantee sample diversity, limiting their performance. We propose a mode-guided diffusion model leveraging a pre-trained diffusion model without the need to fine-tune with distillation losses. Our approach addresses dataset diversity in three stages: Mode Discovery to identify distinct data modes, Mode Guidance to enhance intra-class diversity, and Stop Guidance to mitigate artifacts in synthetic samples that affect performance. Our approach outperforms state-of-the-art methods, achieving accuracy gains of 4.4%, 2.9%, 1.6%, and 1.6% on ImageNette, ImageIDC, ImageNet-100, and ImageNet-1K, respectively. Our method eliminates the need for fine-tuning diffusion models with distillation losses, significantly reducing computational costs. Our code is available on the project webpage: https://jachansantiago.github.io/mode-guided-distillation/
format Preprint
id arxiv_https___arxiv_org_abs_2505_18963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MGD$^3$: Mode-Guided Dataset Distillation using Diffusion Models
Chan-Santiago, Jeffrey A.
Tirupattur, Praveen
Nayak, Gaurav Kumar
Liu, Gaowen
Shah, Mubarak
Computer Vision and Pattern Recognition
Dataset distillation has emerged as an effective strategy, significantly reducing training costs and facilitating more efficient model deployment. Recent advances have leveraged generative models to distill datasets by capturing the underlying data distribution. Unfortunately, existing methods require model fine-tuning with distillation losses to encourage diversity and representativeness. However, these methods do not guarantee sample diversity, limiting their performance. We propose a mode-guided diffusion model leveraging a pre-trained diffusion model without the need to fine-tune with distillation losses. Our approach addresses dataset diversity in three stages: Mode Discovery to identify distinct data modes, Mode Guidance to enhance intra-class diversity, and Stop Guidance to mitigate artifacts in synthetic samples that affect performance. Our approach outperforms state-of-the-art methods, achieving accuracy gains of 4.4%, 2.9%, 1.6%, and 1.6% on ImageNette, ImageIDC, ImageNet-100, and ImageNet-1K, respectively. Our method eliminates the need for fine-tuning diffusion models with distillation losses, significantly reducing computational costs. Our code is available on the project webpage: https://jachansantiago.github.io/mode-guided-distillation/
title MGD$^3$: Mode-Guided Dataset Distillation using Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18963