Saved in:
Bibliographic Details
Main Authors: Ye, He, Kevin, Rojas, Molei, Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.10971
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908406058582016
author Ye, He
Kevin, Rojas
Molei, Tao
author_facet Ye, He
Kevin, Rojas
Molei, Tao
contents We study masked discrete diffusion models with classifier-free guidance (CFG). Assuming no score error nor discretization error, we derive an explicit solution to the guided reverse dynamics, so that how guidance influences the sampling behavior can be precisely characterized. When the full data distribution is a mixture over classes and the goal is to sample from a specific class, guidance amplifies class-specific regions while suppresses regions shared with other classes. This effect depends on the guidance strength $w$ and induces distinct covariance structures in the sampled distribution. Notably, we observe quantitatively different behaviors in $1$D and $2$D. We also show that for large $w$, the decay rate of the total variation ($\mathrm{TV}$) along the reverse dynamics is double-exponential in $w$ for both $1$D and $2$D. These findings highlight the role of guidance, not just in shaping the output distribution, but also in controlling the dynamics of the sampling trajectory. Our theoretical analysis is supported by experiments that illustrate the geometric effects of guidance and its impact on convergence.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10971
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Exactly Does Guidance Do in Masked Discrete Diffusion Models
Ye, He
Kevin, Rojas
Molei, Tao
Machine Learning
We study masked discrete diffusion models with classifier-free guidance (CFG). Assuming no score error nor discretization error, we derive an explicit solution to the guided reverse dynamics, so that how guidance influences the sampling behavior can be precisely characterized. When the full data distribution is a mixture over classes and the goal is to sample from a specific class, guidance amplifies class-specific regions while suppresses regions shared with other classes. This effect depends on the guidance strength $w$ and induces distinct covariance structures in the sampled distribution. Notably, we observe quantitatively different behaviors in $1$D and $2$D. We also show that for large $w$, the decay rate of the total variation ($\mathrm{TV}$) along the reverse dynamics is double-exponential in $w$ for both $1$D and $2$D. These findings highlight the role of guidance, not just in shaping the output distribution, but also in controlling the dynamics of the sampling trajectory. Our theoretical analysis is supported by experiments that illustrate the geometric effects of guidance and its impact on convergence.
title What Exactly Does Guidance Do in Masked Discrete Diffusion Models
topic Machine Learning
url https://arxiv.org/abs/2506.10971