Saved in:
Bibliographic Details
Main Authors: Gao, Yansong, Sun, Yu
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.12889
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915675235155968
author Gao, Yansong
Sun, Yu
author_facet Gao, Yansong
Sun, Yu
contents Discrete diffusion models (DDMs) are a powerful class of generative models for categorical data, but they typically require many function evaluations for a single sample, making inference expensive. Existing acceleration methods either rely on approximate simulators, such as $τ$-leaping, or on distillation schemes that train new student models and auxiliary networks with proxy objectives. We propose a simple and principled distillation alternative based on \emph{conditional distribution matching}. Our key observation is that the reverse conditional distribution of clean data given a noisy state, $p_{0\mid t}(x_0 \mid x_t)$, admits a Markov decomposition through intermediate times and can be recovered from marginal density ratios and the known forward CTMC kernel. We exploit this structure to define distillation objectives that directly match conditional distributions between a pre-trained teacher and a low-NFE student, both for one-step and few-step samplers.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12889
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distillation of Discrete Diffusion by Exact Conditional Distribution Matching
Gao, Yansong
Sun, Yu
Machine Learning
Discrete diffusion models (DDMs) are a powerful class of generative models for categorical data, but they typically require many function evaluations for a single sample, making inference expensive. Existing acceleration methods either rely on approximate simulators, such as $τ$-leaping, or on distillation schemes that train new student models and auxiliary networks with proxy objectives. We propose a simple and principled distillation alternative based on \emph{conditional distribution matching}. Our key observation is that the reverse conditional distribution of clean data given a noisy state, $p_{0\mid t}(x_0 \mid x_t)$, admits a Markov decomposition through intermediate times and can be recovered from marginal density ratios and the known forward CTMC kernel. We exploit this structure to define distillation objectives that directly match conditional distributions between a pre-trained teacher and a low-NFE student, both for one-step and few-step samplers.
title Distillation of Discrete Diffusion by Exact Conditional Distribution Matching
topic Machine Learning
url https://arxiv.org/abs/2512.12889