Empirical Analysis of Decoding Biases in Masked Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Pengcheng, Liu, Tianming, Liu, Zhenghao, Yan, Yukun, Wang, Shuo, Xiao, Tong, Chen, Zulong, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918280432713728
author Huang, Pengcheng
Liu, Tianming
Liu, Zhenghao
Yan, Yukun
Wang, Shuo
Xiao, Tong
Chen, Zulong
Sun, Maosong
author_facet Huang, Pengcheng
Liu, Tianming
Liu, Zhenghao
Yan, Yukun
Wang, Shuo
Xiao, Tong
Chen, Zulong
Sun, Maosong
contents Masked diffusion models (MDMs), which leverage bidirectional attention and a denoising process, are narrowing the performance gap with autoregressive models (ARMs). However, their internal attention mechanisms remain under-explored. This paper investigates the attention behaviors in MDMs, revealing the phenomenon of Attention Floating. Unlike ARMs, where attention converges to a fixed sink, MDMs exhibit dynamic, dispersed attention anchors that shift across denoising steps and layers. Further analysis reveals its Shallow Structure-Aware, Deep Content-Focused attention mechanism: shallow layers utilize floating tokens to build a global structural framework, while deeper layers allocate more capability toward capturing semantic content. Empirically, this distinctive attention pattern provides a mechanistic explanation for the strong in-context learning capabilities of MDMs, allowing them to double the performance compared to ARMs in knowledge-intensive tasks. All codes are available at https://github.com/NEUIR/Uncode.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empirical Analysis of Decoding Biases in Masked Diffusion Models
Huang, Pengcheng
Liu, Tianming
Liu, Zhenghao
Yan, Yukun
Wang, Shuo
Xiao, Tong
Chen, Zulong
Sun, Maosong
Artificial Intelligence
Computation and Language
Masked diffusion models (MDMs), which leverage bidirectional attention and a denoising process, are narrowing the performance gap with autoregressive models (ARMs). However, their internal attention mechanisms remain under-explored. This paper investigates the attention behaviors in MDMs, revealing the phenomenon of Attention Floating. Unlike ARMs, where attention converges to a fixed sink, MDMs exhibit dynamic, dispersed attention anchors that shift across denoising steps and layers. Further analysis reveals its Shallow Structure-Aware, Deep Content-Focused attention mechanism: shallow layers utilize floating tokens to build a global structural framework, while deeper layers allocate more capability toward capturing semantic content. Empirically, this distinctive attention pattern provides a mechanistic explanation for the strong in-context learning capabilities of MDMs, allowing them to double the performance compared to ARMs in knowledge-intensive tasks. All codes are available at https://github.com/NEUIR/Uncode.
title Empirical Analysis of Decoding Biases in Masked Diffusion Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.13021