KLASS: KL-Guided Fast Inference in Masked Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Seo Hyun, Hong, Sunwoo, Jung, Hojung, Park, Youngrok, Yun, Se-Young
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915838346395648
author Kim, Seo Hyun
Hong, Sunwoo
Jung, Hojung
Park, Youngrok
Yun, Se-Young
author_facet Kim, Seo Hyun
Hong, Sunwoo
Jung, Hojung
Park, Youngrok
Yun, Se-Young
contents Masked diffusion models have demonstrated competitive results on various tasks including language generation. However, due to its iterative refinement process, the inference is often bottlenecked by slow and static sampling speed. To overcome this problem, we introduce `KL-Adaptive Stability Sampling' (KLASS), a fast yet effective sampling method that exploits token-level KL divergence to identify stable, high-confidence predictions. By unmasking multiple tokens in each iteration without any additional model training, our approach speeds up generation significantly while maintaining sample quality. On reasoning benchmarks, KLASS achieves up to $2.78\times$ wall-clock speedups while improving performance over standard greedy decoding, attaining state-of-the-art results among diffusion-based samplers. We further validate KLASS across diverse domains, including text, image, and molecular generation, showing its effectiveness as a broadly applicable sampler across different models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05664
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KLASS: KL-Guided Fast Inference in Masked Diffusion Models
Kim, Seo Hyun
Hong, Sunwoo
Jung, Hojung
Park, Youngrok
Yun, Se-Young
Machine Learning
Masked diffusion models have demonstrated competitive results on various tasks including language generation. However, due to its iterative refinement process, the inference is often bottlenecked by slow and static sampling speed. To overcome this problem, we introduce `KL-Adaptive Stability Sampling' (KLASS), a fast yet effective sampling method that exploits token-level KL divergence to identify stable, high-confidence predictions. By unmasking multiple tokens in each iteration without any additional model training, our approach speeds up generation significantly while maintaining sample quality. On reasoning benchmarks, KLASS achieves up to $2.78\times$ wall-clock speedups while improving performance over standard greedy decoding, attaining state-of-the-art results among diffusion-based samplers. We further validate KLASS across diverse domains, including text, image, and molecular generation, showing its effectiveness as a broadly applicable sampler across different models.
title KLASS: KL-Guided Fast Inference in Masked Diffusion Models
topic Machine Learning
url https://arxiv.org/abs/2511.05664