DiLaDiff: Distilled Latent-Augmented Diffusion for Language Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911709047816192 |
|---|---|
| author | Lemercier, Jean-Marie Geffner, Tomas Kreis, Karsten Mardani, Morteza Vahdat, Arash Jukić, Ante |
| author_facet | Lemercier, Jean-Marie Geffner, Tomas Kreis, Karsten Mardani, Morteza Vahdat, Arash Jukić, Ante |
| contents | Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_23605 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DiLaDiff: Distilled Latent-Augmented Diffusion for Language Modeling Lemercier, Jean-Marie Geffner, Tomas Kreis, Karsten Mardani, Morteza Vahdat, Arash Jukić, Ante Machine Learning Artificial Intelligence Computation and Language Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding. |
| title | DiLaDiff: Distilled Latent-Augmented Diffusion for Language Modeling |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2605.23605 |