Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Tianyu, Xue, Shuchen, Feng, Zijin, Hu, Tianyang, Sun, Jiacheng, Li, Zhenguo, Zhang, Cheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910126634434560
author Xie, Tianyu
Xue, Shuchen
Feng, Zijin
Hu, Tianyang
Sun, Jiacheng
Li, Zhenguo
Zhang, Cheng
author_facet Xie, Tianyu
Xue, Shuchen
Feng, Zijin
Hu, Tianyang
Sun, Jiacheng
Li, Zhenguo
Zhang, Cheng
contents Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a compelling trade-off between quality and generation speed. MDMs denoise by progressively unmasking multiple dimensions from an all-masked input, but their performance can degrade when using few denoising steps due to limited modeling of inter-dimensional dependencies. In this paper, we propose Variational Autoencoding Discrete Diffusion (VADD), a novel framework that enhances discrete diffusion with latent variable modeling to implicitly capture correlations among dimensions. By introducing an auxiliary recognition model, VADD enables stable training via variational lower bounds maximization and amortized inference over the training set. Our approach retains the efficiency of traditional MDMs while significantly improving sample quality, especially when the number of denoising steps is small. Empirical results on 2D toy data, pixel-level image generation, and text generation demonstrate that VADD consistently outperforms MDM baselines in sample quality with few denoising steps.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17384
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling
Xie, Tianyu
Xue, Shuchen
Feng, Zijin
Hu, Tianyang
Sun, Jiacheng
Li, Zhenguo
Zhang, Cheng
Machine Learning
Computer Vision and Pattern Recognition
Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a compelling trade-off between quality and generation speed. MDMs denoise by progressively unmasking multiple dimensions from an all-masked input, but their performance can degrade when using few denoising steps due to limited modeling of inter-dimensional dependencies. In this paper, we propose Variational Autoencoding Discrete Diffusion (VADD), a novel framework that enhances discrete diffusion with latent variable modeling to implicitly capture correlations among dimensions. By introducing an auxiliary recognition model, VADD enables stable training via variational lower bounds maximization and amortized inference over the training set. Our approach retains the efficiency of traditional MDMs while significantly improving sample quality, especially when the number of denoising steps is small. Empirical results on 2D toy data, pixel-level image generation, and text generation demonstrate that VADD consistently outperforms MDM baselines in sample quality with few denoising steps.
title Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17384