ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Jia-Nan, Guan, Jian, Wu, Wei, Li, Chongxuan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908867722477568
author Li, Jia-Nan
Guan, Jian
Wu, Wei
Li, Chongxuan
author_facet Li, Jia-Nan
Guan, Jian
Wu, Wei
Li, Chongxuan
contents Autoregressive models (ARMs) are hindered by slow sequential inference. While masked diffusion models (MDMs) offer a parallel alternative, they suffer from critical drawbacks: high computational overhead from precluding Key-Value (KV) caching, and incoherent generation arising from learning dependencies over an intractable space of token combinations. To address these limitations, we introduce \textsc{ReFusion}, a novel masked diffusion model that integrates sequence reorganization into the causal attention framework. By elevating parallel decoding from the token level to a higher slot level, \textsc{ReFusion} interleaves inter-slot diffusion-based selection with intra-slot autoregressive infilling, while reordering newly generated slots ahead of the remaining masks after each iteration. Consequently, this design simultaneously unlocks full KV cache reuse and reduces learning complexity from an intractable token combination space to a manageable slot-level permutation space. Extensive experiments on seven diverse benchmarks show that \textsc{ReFusion} not only overwhelmingly surpasses prior MDMs with a 34\% performance gain and an over 18$\times$ speedup on average, but also bridges the performance gap to strong ARMs while maintaining a 2.33$\times$ average speedup.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
Li, Jia-Nan
Guan, Jian
Wu, Wei
Li, Chongxuan
Computation and Language
Artificial Intelligence
Machine Learning
Autoregressive models (ARMs) are hindered by slow sequential inference. While masked diffusion models (MDMs) offer a parallel alternative, they suffer from critical drawbacks: high computational overhead from precluding Key-Value (KV) caching, and incoherent generation arising from learning dependencies over an intractable space of token combinations. To address these limitations, we introduce \textsc{ReFusion}, a novel masked diffusion model that integrates sequence reorganization into the causal attention framework. By elevating parallel decoding from the token level to a higher slot level, \textsc{ReFusion} interleaves inter-slot diffusion-based selection with intra-slot autoregressive infilling, while reordering newly generated slots ahead of the remaining masks after each iteration. Consequently, this design simultaneously unlocks full KV cache reuse and reduces learning complexity from an intractable token combination space to a manageable slot-level permutation space. Extensive experiments on seven diverse benchmarks show that \textsc{ReFusion} not only overwhelmingly surpasses prior MDMs with a 34\% performance gain and an over 18$\times$ speedup on average, but also bridges the performance gap to strong ARMs while maintaining a 2.33$\times$ average speedup.
title ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.13586