Bringing Stability to Diffusion: Decomposing and Reducing Variance of Training Masked Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jia, Mengni, Zhou, Mengyu, Liu, Yihao, Jiang, Xiaoxi, Jiang, Guanjun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917518493351936
author Jia, Mengni
Zhou, Mengyu
Liu, Yihao
Jiang, Xiaoxi
Jiang, Guanjun
author_facet Jia, Mengni
Zhou, Mengyu
Liu, Yihao
Jiang, Xiaoxi
Jiang, Guanjun
contents Masked diffusion models (MDMs) are a promising alternative to autoregressive models (ARMs), but they suffer from inherently much higher training variance. High variance leads to noisier gradient estimates and unstable optimization, so even equally strong pretrained MDMs and ARMs that are competitive at initialization often diverge after task-specific training, with MDMs falling far behind. There has been no theoretical explanation or systematic solution. We derive the first decomposition of MDM training variance into three sources: (A) masking pattern noise, (B) masking rate noise, and (C) data noise, while ARMs are only affected by (C). This explains the fundamental training gap. Building on this foundation, we design six variance-reduction methods, including two core methods: (1) P-POTS, a Pareto-optimal t sampler that minimizes training variance by sampling harder t values more often with appropriately smaller update steps, and (2) MIRROR, which uses negatively correlated samples to reduce (A). Experiments show that compared to standard MDM training, our methods improve accuracy by 7-8% on complex reasoning tasks, while simultaneously reducing run-to-run variability to near ARM levels, substantially narrowing the gap with strong ARM baselines; in most settings, even the best baseline runs remain below the worst run of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18159
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bringing Stability to Diffusion: Decomposing and Reducing Variance of Training Masked Diffusion Models
Jia, Mengni
Zhou, Mengyu
Liu, Yihao
Jiang, Xiaoxi
Jiang, Guanjun
Machine Learning
Masked diffusion models (MDMs) are a promising alternative to autoregressive models (ARMs), but they suffer from inherently much higher training variance. High variance leads to noisier gradient estimates and unstable optimization, so even equally strong pretrained MDMs and ARMs that are competitive at initialization often diverge after task-specific training, with MDMs falling far behind. There has been no theoretical explanation or systematic solution. We derive the first decomposition of MDM training variance into three sources: (A) masking pattern noise, (B) masking rate noise, and (C) data noise, while ARMs are only affected by (C). This explains the fundamental training gap. Building on this foundation, we design six variance-reduction methods, including two core methods: (1) P-POTS, a Pareto-optimal t sampler that minimizes training variance by sampling harder t values more often with appropriately smaller update steps, and (2) MIRROR, which uses negatively correlated samples to reduce (A). Experiments show that compared to standard MDM training, our methods improve accuracy by 7-8% on complex reasoning tasks, while simultaneously reducing run-to-run variability to near ARM levels, substantially narrowing the gap with strong ARM baselines; in most settings, even the best baseline runs remain below the worst run of our method.
title Bringing Stability to Diffusion: Decomposing and Reducing Variance of Training Masked Diffusion Models
topic Machine Learning
url https://arxiv.org/abs/2511.18159