Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Seo, Yeongbin, Lee, Dongha, Kim, Jaehyung, Yeo, Jinyoung
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917038391296000
author Seo, Yeongbin
Lee, Dongha
Kim, Jaehyung
Yeo, Jinyoung
author_facet Seo, Yeongbin
Lee, Dongha
Kim, Jaehyung
Yeo, Jinyoung
contents Autoregressive (AR) language models generate text one token at a time, which limits their inference speed. Diffusion-based language models offer a promising alternative, as they can decode multiple tokens in parallel. However, we identify a key bottleneck in current diffusion LMs: the long decoding-window problem, where tokens generated far from the input context often become irrelevant or repetitive. Previous solutions like semi-autoregressive address this issue by splitting windows into blocks (sacrificing bidirectionality), but we find that this also leads to time-interval expansion problem, sacrificing the speed. Therefore, semi-AR eliminates the main advantages of diffusion models. To overcome this, we propose Convolutional decoding (Conv), a normalization-based method that narrows the decoding window without hard segmentation, leading to better fluency and flexibility. Additionally, we introduce Rejecting Rule-based Fine-Tuning (R2FT), a post-hoc training scheme that better aligns tokens at positions far from context. Our methods achieve state-of-the-art results on open-ended generation benchmarks (e.g., AlpacaEval) among diffusion LM baselines, with significantly lower step size than previous works, demonstrating both speed and quality improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
Seo, Yeongbin
Lee, Dongha
Kim, Jaehyung
Yeo, Jinyoung
Computation and Language
Artificial Intelligence
Machine Learning
68T50
I.2.7
Autoregressive (AR) language models generate text one token at a time, which limits their inference speed. Diffusion-based language models offer a promising alternative, as they can decode multiple tokens in parallel. However, we identify a key bottleneck in current diffusion LMs: the long decoding-window problem, where tokens generated far from the input context often become irrelevant or repetitive. Previous solutions like semi-autoregressive address this issue by splitting windows into blocks (sacrificing bidirectionality), but we find that this also leads to time-interval expansion problem, sacrificing the speed. Therefore, semi-AR eliminates the main advantages of diffusion models. To overcome this, we propose Convolutional decoding (Conv), a normalization-based method that narrows the decoding window without hard segmentation, leading to better fluency and flexibility. Additionally, we introduce Rejecting Rule-based Fine-Tuning (R2FT), a post-hoc training scheme that better aligns tokens at positions far from context. Our methods achieve state-of-the-art results on open-ended generation benchmarks (e.g., AlpacaEval) among diffusion LM baselines, with significantly lower step size than previous works, demonstrating both speed and quality improvements.
title Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
topic Computation and Language
Artificial Intelligence
Machine Learning
68T50
I.2.7
url https://arxiv.org/abs/2509.15188