Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Shuchen, Xie, Tianyu, Hu, Tianyang, Feng, Zijin, Sun, Jiacheng, Kawaguchi, Kenji, Li, Zhenguo, Ma, Zhi-Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913911485235200
author Xue, Shuchen
Xie, Tianyu
Hu, Tianyang
Feng, Zijin
Sun, Jiacheng
Kawaguchi, Kenji
Li, Zhenguo
Ma, Zhi-Ming
author_facet Xue, Shuchen
Xie, Tianyu
Hu, Tianyang
Feng, Zijin
Sun, Jiacheng
Kawaguchi, Kenji
Li, Zhenguo
Ma, Zhi-Ming
contents Large language models (LLMs) predominantly use autoregressive (AR) approaches, but masked diffusion models (MDMs) are emerging as viable alternatives. A key challenge in comparing AR and MDM paradigms is their typical architectural difference: AR models are often decoder-only, while MDMs have largely been encoder-only. This practice of changing both the modeling paradigm and architecture simultaneously makes direct comparisons unfair, as it's hard to distinguish whether observed differences stem from the paradigm itself or the architectural shift. This research evaluates MDMs within a decoder-only framework to: (1) equitably compare MDM (as Any-Order AR, or AO-AR) and standard AR paradigms. Our investigation suggests that the standard AO-AR objective, which averages over all token permutations, may benefit from refinement, as many permutations appear less informative compared to the language's inherent left-to-right structure. (2) Investigate architectural influences (decoder-only vs. encoder-only) within MDMs. We demonstrate that while encoder-only MDMs model a simpler conditional probability space, decoder-only MDMs can achieve dramatic generation speedups ($\sim25\times$) and comparable perplexity with temperature annealing despite modeling a vastly larger space, highlighting key trade-offs. This work thus decouples core paradigm differences from architectural influences, offering insights for future model design. Code is available at https://github.com/scxue/AO-GPT-MDM.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19935
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
Xue, Shuchen
Xie, Tianyu
Hu, Tianyang
Feng, Zijin
Sun, Jiacheng
Kawaguchi, Kenji
Li, Zhenguo
Ma, Zhi-Ming
Machine Learning
Computer Vision and Pattern Recognition
Large language models (LLMs) predominantly use autoregressive (AR) approaches, but masked diffusion models (MDMs) are emerging as viable alternatives. A key challenge in comparing AR and MDM paradigms is their typical architectural difference: AR models are often decoder-only, while MDMs have largely been encoder-only. This practice of changing both the modeling paradigm and architecture simultaneously makes direct comparisons unfair, as it's hard to distinguish whether observed differences stem from the paradigm itself or the architectural shift. This research evaluates MDMs within a decoder-only framework to: (1) equitably compare MDM (as Any-Order AR, or AO-AR) and standard AR paradigms. Our investigation suggests that the standard AO-AR objective, which averages over all token permutations, may benefit from refinement, as many permutations appear less informative compared to the language's inherent left-to-right structure. (2) Investigate architectural influences (decoder-only vs. encoder-only) within MDMs. We demonstrate that while encoder-only MDMs model a simpler conditional probability space, decoder-only MDMs can achieve dramatic generation speedups ($\sim25\times$) and comparable perplexity with temperature annealing despite modeling a vastly larger space, highlighting key trade-offs. This work thus decouples core paradigm differences from architectural influences, offering insights for future model design. Code is available at https://github.com/scxue/AO-GPT-MDM.
title Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.19935