DiLaDiff: Distilled Latent-Augmented Diffusion for Language Modeling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lemercier, Jean-Marie, Geffner, Tomas, Kreis, Karsten, Mardani, Morteza, Vahdat, Arash, Jukić, Ante
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911709047816192
author Lemercier, Jean-Marie
Geffner, Tomas
Kreis, Karsten
Mardani, Morteza
Vahdat, Arash
Jukić, Ante
author_facet Lemercier, Jean-Marie
Geffner, Tomas
Kreis, Karsten
Mardani, Morteza
Vahdat, Arash
Jukić, Ante
contents Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23605
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DiLaDiff: Distilled Latent-Augmented Diffusion for Language Modeling
Lemercier, Jean-Marie
Geffner, Tomas
Kreis, Karsten
Mardani, Morteza
Vahdat, Arash
Jukić, Ante
Machine Learning
Artificial Intelligence
Computation and Language
Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding.
title DiLaDiff: Distilled Latent-Augmented Diffusion for Language Modeling
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.23605