Saved in:
Bibliographic Details
Main Authors: Wu, Junyi, Zhao, Tianchen, Zhang, Shaoqiu, Zhang, Linfeng, Dai, Guohao, Wang, Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.18165
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914577586847744
author Wu, Junyi
Zhao, Tianchen
Zhang, Shaoqiu
Zhang, Linfeng
Dai, Guohao
Wang, Yu
author_facet Wu, Junyi
Zhao, Tianchen
Zhang, Shaoqiu
Zhang, Linfeng
Dai, Guohao
Wang, Yu
contents Unlike autoregressive models, which generate one token at a time, dLLMs denoise a chunk of [MASK] tokens jointly and sample one or more tokens per step; despite enabling parallel decoding, this process incurs substantial computational cost due to the large chunk size of masked tokens. We observe that much of this cost is spent on repeatedly processing the preceding context and many [MASK] tokens with the same feature representations, indicating considerable computational redundancy. In this work, we revisit dLLM's redundancy from the perspective of [MASK] tokens. Through systematic analysis, we verify the redundancy of [MASK] tokens while revealing their critical role in providing structural information. Guided by these findings, we propose position-preserving [MASK] token compression and terminal-aware augmentation. By compressing redundant [MASK] computation, this approach accelerates decoding and further provides a natural extension toward context-folding-like long-context scaling under limited input-length constraints for full-sequence dLLMs such as LLaDA-8B-Instruct and LLaDA-1.5. Moreover, for block dLLMs such as LLaDA2.0-mini, it augments the context with a protected terminal [MASK] token to enhance generation quality with negligible overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18165
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Elastic-dLLM: Position Preserving Context Compression and Augmentation of Diffusion LLMs
Wu, Junyi
Zhao, Tianchen
Zhang, Shaoqiu
Zhang, Linfeng
Dai, Guohao
Wang, Yu
Machine Learning
Unlike autoregressive models, which generate one token at a time, dLLMs denoise a chunk of [MASK] tokens jointly and sample one or more tokens per step; despite enabling parallel decoding, this process incurs substantial computational cost due to the large chunk size of masked tokens. We observe that much of this cost is spent on repeatedly processing the preceding context and many [MASK] tokens with the same feature representations, indicating considerable computational redundancy. In this work, we revisit dLLM's redundancy from the perspective of [MASK] tokens. Through systematic analysis, we verify the redundancy of [MASK] tokens while revealing their critical role in providing structural information. Guided by these findings, we propose position-preserving [MASK] token compression and terminal-aware augmentation. By compressing redundant [MASK] computation, this approach accelerates decoding and further provides a natural extension toward context-folding-like long-context scaling under limited input-length constraints for full-sequence dLLMs such as LLaDA-8B-Instruct and LLaDA-1.5. Moreover, for block dLLMs such as LLaDA2.0-mini, it augments the context with a protected terminal [MASK] token to enhance generation quality with negligible overhead.
title Elastic-dLLM: Position Preserving Context Compression and Augmentation of Diffusion LLMs
topic Machine Learning
url https://arxiv.org/abs/2605.18165