CDLM: Consistency Diffusion Language Models For Faster Sampling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Minseo, Xu, Chenfeng, Hooper, Coleman, Singh, Harman, Athiwaratkun, Ben, Zhang, Ce, Keutzer, Kurt, Gholami, Amir
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912914639683584
author Kim, Minseo
Xu, Chenfeng
Hooper, Coleman
Singh, Harman
Athiwaratkun, Ben
Zhang, Ce
Keutzer, Kurt
Gholami, Amir
author_facet Kim, Minseo
Xu, Chenfeng
Hooper, Coleman
Singh, Harman
Athiwaratkun, Ben
Zhang, Ce
Keutzer, Kurt
Gholami, Amir
contents Diffusion Language Models (DLMs) offer a promising parallel generation paradigm but suffer from slow inference due to numerous refinement steps and the inability to use standard KV caching. We introduce CDLM (Consistency Diffusion Language Models), a training-based acceleration method that simultaneously tackles both bottlenecks. CDLM integrates consistency modeling to drastically reduce the number of required sampling steps by enabling multi-token finalization. Furthermore, we enforce a block-wise causal attention mask during fine-tuning, making the model fully compatible with KV caching. Experiments show CDLM achieves 3.6x-14.5x lower latency while maintaining competitive accuracy on math and coding tasks. The full training and evaluation code is available at https://github.com/SqueezeAILab/CDLM.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CDLM: Consistency Diffusion Language Models For Faster Sampling
Kim, Minseo
Xu, Chenfeng
Hooper, Coleman
Singh, Harman
Athiwaratkun, Ben
Zhang, Ce
Keutzer, Kurt
Gholami, Amir
Machine Learning
Computation and Language
Diffusion Language Models (DLMs) offer a promising parallel generation paradigm but suffer from slow inference due to numerous refinement steps and the inability to use standard KV caching. We introduce CDLM (Consistency Diffusion Language Models), a training-based acceleration method that simultaneously tackles both bottlenecks. CDLM integrates consistency modeling to drastically reduce the number of required sampling steps by enabling multi-token finalization. Furthermore, we enforce a block-wise causal attention mask during fine-tuning, making the model fully compatible with KV caching. Experiments show CDLM achieves 3.6x-14.5x lower latency while maintaining competitive accuracy on math and coding tasks. The full training and evaluation code is available at https://github.com/SqueezeAILab/CDLM.
title CDLM: Consistency Diffusion Language Models For Faster Sampling
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2511.19269