Focus on the Core: Empowering Diffusion Large Language Models by Self-Contrast

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Feng, Jinyuan, Yu, Xin, Chen, Yiqun, Wei, Xiaochi, Gao, Yan, Wu, Yi, Hu, Yao, Pu, Zhiqiang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915974197805056
author Feng, Jinyuan
Yu, Xin
Chen, Yiqun
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Pu, Zhiqiang
author_facet Feng, Jinyuan
Yu, Xin
Chen, Yiqun
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Pu, Zhiqiang
contents The iterative denoising paradigm of Diffusion Large Language Models (DLMs) endows them with a distinct advantage in global context modeling. However, current decoding strategies fail to leverage this capability, typically exhibiting a local preference that overlooks the heterogeneous information density within the context, ultimately degrading generation quality. To address this limitation, we systematically investigate high-information-density (HD) tokens and present two key findings: (1) explicitly conditioning on HD tokens substantially improves output quality; and (2) HD tokens exhibit an early-decoding tendency, converging earlier than surrounding tokens. Motivated by these findings, we propose Focus on the Core \textbf{(FoCore)}, a training-free decoding strategy that utilizes HD tokens in a self-contrast manner, wherein HD tokens are temporarily remasked as negative samples, to guide generation. We further introduce FoCore\_Accelerate \textbf{(FoCore\_A)}, an efficient variant that, upon detecting HD token convergence, performs parallel decoding over stable candidates within a local context window, substantially accelerating generation. Extensive experiments on math, code and logical reasoning benchmarks demonstrate that FoCore consistently improves generation quality and efficiency across both LLaDA and Dream backbones. For instance, on HumanEval, FoCore improves pass@1 from 39.02 to 42.68 over standard Classifier-Free Guidance, while FoCore-A reduces the number of decoding steps by 2.07x and per-sample latency from 20.76s to 8.64s (-58.4\%).
format Preprint
id arxiv_https___arxiv_org_abs_2605_01373
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Focus on the Core: Empowering Diffusion Large Language Models by Self-Contrast
Feng, Jinyuan
Yu, Xin
Chen, Yiqun
Wei, Xiaochi
Gao, Yan
Wu, Yi
Hu, Yao
Pu, Zhiqiang
Computation and Language
Artificial Intelligence
The iterative denoising paradigm of Diffusion Large Language Models (DLMs) endows them with a distinct advantage in global context modeling. However, current decoding strategies fail to leverage this capability, typically exhibiting a local preference that overlooks the heterogeneous information density within the context, ultimately degrading generation quality. To address this limitation, we systematically investigate high-information-density (HD) tokens and present two key findings: (1) explicitly conditioning on HD tokens substantially improves output quality; and (2) HD tokens exhibit an early-decoding tendency, converging earlier than surrounding tokens. Motivated by these findings, we propose Focus on the Core \textbf{(FoCore)}, a training-free decoding strategy that utilizes HD tokens in a self-contrast manner, wherein HD tokens are temporarily remasked as negative samples, to guide generation. We further introduce FoCore\_Accelerate \textbf{(FoCore\_A)}, an efficient variant that, upon detecting HD token convergence, performs parallel decoding over stable candidates within a local context window, substantially accelerating generation. Extensive experiments on math, code and logical reasoning benchmarks demonstrate that FoCore consistently improves generation quality and efficiency across both LLaDA and Dream backbones. For instance, on HumanEval, FoCore improves pass@1 from 39.02 to 42.68 over standard Classifier-Free Guidance, while FoCore-A reduces the number of decoding steps by 2.07x and per-sample latency from 20.76s to 8.64s (-58.4\%).
title Focus on the Core: Empowering Diffusion Large Language Models by Self-Contrast
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.01373