Differences in Text Generated by Diffusion and Autoregressive Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zeyang, Liang, Chengwei, Chen, Xingyan, Gu, Meiqi, Luo, Minrui, Zhang, Jingzhao, He, Tianxing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914560639762432
author Zhang, Zeyang
Liang, Chengwei
Chen, Xingyan
Gu, Meiqi
Luo, Minrui
Zhang, Jingzhao
He, Tianxing
author_facet Zhang, Zeyang
Liang, Chengwei
Chen, Xingyan
Gu, Meiqi
Luo, Minrui
Zhang, Jingzhao
He, Tianxing
contents Diffusion language models (DLMs) are promising alternatives to autoregressive language models (ARMs), yet the intrinsic differences in their generated text remain underexplored. We first find empirically that off-the-shelf DLMs exhibit lower $n$-gram entropy, higher semantic coherence, and higher semantic diversity. To understand the cause, we conduct controlled experiments that decouple the effects of training objectives and decoding algorithms. Results suggest that the DLM training objective contributes to the increases in semantic coherence and semantic diversity, but has a minor influence on entropy. These differences are primarily driven by the bidirectional context; other components in the training objective, such as input masking, label masking, and the weighting function, have a much weaker influence. Further, our experiments demonstrate that the reduction in entropy stems from DLMs' decoding algorithms, particularly confidence-based remasking strategies. We provide a theoretical understanding for this entropy reduction phenomenon. Together, our work uncovers key mechanisms underlying the differences between DLMs and ARMs in text generation, and informs future design of training objectives and decoding algorithms in DLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12522
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Differences in Text Generated by Diffusion and Autoregressive Language Models
Zhang, Zeyang
Liang, Chengwei
Chen, Xingyan
Gu, Meiqi
Luo, Minrui
Zhang, Jingzhao
He, Tianxing
Computation and Language
Artificial Intelligence
Diffusion language models (DLMs) are promising alternatives to autoregressive language models (ARMs), yet the intrinsic differences in their generated text remain underexplored. We first find empirically that off-the-shelf DLMs exhibit lower $n$-gram entropy, higher semantic coherence, and higher semantic diversity. To understand the cause, we conduct controlled experiments that decouple the effects of training objectives and decoding algorithms. Results suggest that the DLM training objective contributes to the increases in semantic coherence and semantic diversity, but has a minor influence on entropy. These differences are primarily driven by the bidirectional context; other components in the training objective, such as input masking, label masking, and the weighting function, have a much weaker influence. Further, our experiments demonstrate that the reduction in entropy stems from DLMs' decoding algorithms, particularly confidence-based remasking strategies. We provide a theoretical understanding for this entropy reduction phenomenon. Together, our work uncovers key mechanisms underlying the differences between DLMs and ARMs in text generation, and informs future design of training objectives and decoding algorithms in DLMs.
title Differences in Text Generated by Diffusion and Autoregressive Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.12522