Simplified and Generalized Masked Diffusion for Discrete Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Jiaxin, Han, Kehang, Wang, Zhe, Doucet, Arnaud, Titsias, Michalis K.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913652759592960
author Shi, Jiaxin
Han, Kehang
Wang, Zhe
Doucet, Arnaud
Titsias, Michalis K.
author_facet Shi, Jiaxin
Han, Kehang
Wang, Zhe
Doucet, Arnaud
Titsias, Michalis K.
contents Masked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has been hindered by unnecessarily complex model formulations and unclear relationships between different perspectives, leading to suboptimal parameterization, training objectives, and ad hoc adjustments to counteract these issues. In this work, we aim to provide a simple and general framework that unlocks the full potential of masked diffusion models. We show that the continuous-time variational objective of masked diffusion models is a simple weighted integral of cross-entropy losses. Our framework also enables training generalized masked diffusion models with state-dependent masking schedules. When evaluated by perplexity, our models trained on OpenWebText surpass prior diffusion language models at GPT-2 scale and demonstrate superior performance on 4 out of 5 zero-shot language modeling tasks. Furthermore, our models vastly outperform previous discrete diffusion models on pixel-level image modeling, achieving 2.75 (CIFAR-10) and 3.40 (ImageNet 64x64) bits per dimension that are better than autoregressive models of similar sizes. Our code is available at https://github.com/google-deepmind/md4.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04329
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Simplified and Generalized Masked Diffusion for Discrete Data
Shi, Jiaxin
Han, Kehang
Wang, Zhe
Doucet, Arnaud
Titsias, Michalis K.
Machine Learning
Masked (or absorbing) diffusion is actively explored as an alternative to autoregressive models for generative modeling of discrete data. However, existing work in this area has been hindered by unnecessarily complex model formulations and unclear relationships between different perspectives, leading to suboptimal parameterization, training objectives, and ad hoc adjustments to counteract these issues. In this work, we aim to provide a simple and general framework that unlocks the full potential of masked diffusion models. We show that the continuous-time variational objective of masked diffusion models is a simple weighted integral of cross-entropy losses. Our framework also enables training generalized masked diffusion models with state-dependent masking schedules. When evaluated by perplexity, our models trained on OpenWebText surpass prior diffusion language models at GPT-2 scale and demonstrate superior performance on 4 out of 5 zero-shot language modeling tasks. Furthermore, our models vastly outperform previous discrete diffusion models on pixel-level image modeling, achieving 2.75 (CIFAR-10) and 3.40 (ImageNet 64x64) bits per dimension that are better than autoregressive models of similar sizes. Our code is available at https://github.com/google-deepmind/md4.
title Simplified and Generalized Masked Diffusion for Discrete Data
topic Machine Learning
url https://arxiv.org/abs/2406.04329