From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Yuchuan, Liang, Yuchen, Zhang, Shuo, Shu, Yingte, Yang, Guangwen, He, Wei, Fang, Sibo, Guo, Tianyu, Han, Kai, Xu, Chao, Chen, Hanting, Chen, Xinghao, Wang, Yunhe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911410343116800
author Tian, Yuchuan
Liang, Yuchen
Zhang, Shuo
Shu, Yingte
Yang, Guangwen
He, Wei
Fang, Sibo
Guo, Tianyu
Han, Kai
Xu, Chao
Chen, Hanting
Chen, Xinghao
Wang, Yunhe
author_facet Tian, Yuchuan
Liang, Yuchen
Zhang, Shuo
Shu, Yingte
Yang, Guangwen
He, Wei
Fang, Sibo
Guo, Tianyu
Han, Kai
Xu, Chao
Chen, Hanting
Chen, Xinghao
Wang, Yunhe
contents Diffusion Language Models (DLMs) enable fast generation, yet training large DLMs from scratch is costly. As a practical shortcut, adapting off-the-shelf Auto-Regressive (AR) model weights into a DLM could quickly equip the DLM with strong long-context generation capabilies. Prior "adaptation" attempts either modify logits or randomly grow attention masks to Full-Sequence diffusion, or simply transplant AR weights into a Block-Diffusion recipe, leaving two key questions unaddressed: where is the final destination of adaptation, and how to adapt better? For manifold benefits, we reframe the whole AR-to-DLM adaptation under the Block-Diffusion paradigm, transitioning from block size 1 to the final Block-Diffusion state. Concretely, the principled pathway of adaptation is designed as follows: we keep a context-causal path where causal attention is kept in the prefix, an efficient parallel adaptation procedure where an AR guidance is maintained, and gradual increment of the generation block size for a smoother transition. Built on these components, the adaptation is proved competitive on various models at different scales. With better adaptation, we propose NBDiff-7B that could inherit the long-context modeling and reasoning capabilities, and achieve state-of-the-art performance among the 7B-class DLMs. Codes: https://github.com/YuchuanTian/NBDiff.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06776
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs
Tian, Yuchuan
Liang, Yuchen
Zhang, Shuo
Shu, Yingte
Yang, Guangwen
He, Wei
Fang, Sibo
Guo, Tianyu
Han, Kai
Xu, Chao
Chen, Hanting
Chen, Xinghao
Wang, Yunhe
Computation and Language
Artificial Intelligence
Diffusion Language Models (DLMs) enable fast generation, yet training large DLMs from scratch is costly. As a practical shortcut, adapting off-the-shelf Auto-Regressive (AR) model weights into a DLM could quickly equip the DLM with strong long-context generation capabilies. Prior "adaptation" attempts either modify logits or randomly grow attention masks to Full-Sequence diffusion, or simply transplant AR weights into a Block-Diffusion recipe, leaving two key questions unaddressed: where is the final destination of adaptation, and how to adapt better? For manifold benefits, we reframe the whole AR-to-DLM adaptation under the Block-Diffusion paradigm, transitioning from block size 1 to the final Block-Diffusion state. Concretely, the principled pathway of adaptation is designed as follows: we keep a context-causal path where causal attention is kept in the prefix, an efficient parallel adaptation procedure where an AR guidance is maintained, and gradual increment of the generation block size for a smoother transition. Built on these components, the adaptation is proved competitive on various models at different scales. With better adaptation, we propose NBDiff-7B that could inherit the long-context modeling and reasoning capabilities, and achieve state-of-the-art performance among the 7B-class DLMs. Codes: https://github.com/YuchuanTian/NBDiff.
title From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.06776