Generative Distribution Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Jiequan, Zhu, Beier, Xu, Qingshan, Xu, Xiaogang, Chen, Pengguang, Qi, Xiaojuan, Yu, Bei, Zhang, Hanwang, Hong, Richang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912492681166848
author Cui, Jiequan
Zhu, Beier
Xu, Qingshan
Xu, Xiaogang
Chen, Pengguang
Qi, Xiaojuan
Yu, Bei
Zhang, Hanwang
Hong, Richang
author_facet Cui, Jiequan
Zhu, Beier
Xu, Qingshan
Xu, Xiaogang
Chen, Pengguang
Qi, Xiaojuan
Yu, Bei
Zhang, Hanwang
Hong, Richang
contents In this paper, we formulate the knowledge distillation (KD) as a conditional generative problem and propose the \textit{Generative Distribution Distillation (GenDD)} framework. A naive \textit{GenDD} baseline encounters two major challenges: the curse of high-dimensional optimization and the lack of semantic supervision from labels. To address these issues, we introduce a \textit{Split Tokenization} strategy, achieving stable and effective unsupervised KD. Additionally, we develop the \textit{Distribution Contraction} technique to integrate label supervision into the reconstruction objective. Our theoretical proof demonstrates that \textit{GenDD} with \textit{Distribution Contraction} serves as a gradient-level surrogate for multi-task learning, realizing efficient supervised training without explicit classification loss on multi-step sampling image representations. To evaluate the effectiveness of our method, we conduct experiments on balanced, imbalanced, and unlabeled data. Experimental results show that \textit{GenDD} performs competitively in the unsupervised setting, significantly surpassing KL baseline by \textbf{16.29\%} on ImageNet validation set. With label supervision, our ResNet-50 achieves \textbf{82.28\%} top-1 accuracy on ImageNet in 600 epochs training, establishing a new state-of-the-art.
format Preprint
id arxiv_https___arxiv_org_abs_2507_14503
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Distribution Distillation
Cui, Jiequan
Zhu, Beier
Xu, Qingshan
Xu, Xiaogang
Chen, Pengguang
Qi, Xiaojuan
Yu, Bei
Zhang, Hanwang
Hong, Richang
Machine Learning
Computer Vision and Pattern Recognition
In this paper, we formulate the knowledge distillation (KD) as a conditional generative problem and propose the \textit{Generative Distribution Distillation (GenDD)} framework. A naive \textit{GenDD} baseline encounters two major challenges: the curse of high-dimensional optimization and the lack of semantic supervision from labels. To address these issues, we introduce a \textit{Split Tokenization} strategy, achieving stable and effective unsupervised KD. Additionally, we develop the \textit{Distribution Contraction} technique to integrate label supervision into the reconstruction objective. Our theoretical proof demonstrates that \textit{GenDD} with \textit{Distribution Contraction} serves as a gradient-level surrogate for multi-task learning, realizing efficient supervised training without explicit classification loss on multi-step sampling image representations. To evaluate the effectiveness of our method, we conduct experiments on balanced, imbalanced, and unlabeled data. Experimental results show that \textit{GenDD} performs competitively in the unsupervised setting, significantly surpassing KL baseline by \textbf{16.29\%} on ImageNet validation set. With label supervision, our ResNet-50 achieves \textbf{82.28\%} top-1 accuracy on ImageNet in 600 epochs training, establishing a new state-of-the-art.
title Generative Distribution Distillation
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.14503