Discrete Stochastic Localization for Non-autoregressive Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Yunshu, Cheng, Jiayi, Yu, Longxuan, Thakuria, Partha, Brekelmans, Rob, Papalexakis, Evangelos E., Steeg, Greg Ver
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914584394203136
author Wu, Yunshu
Cheng, Jiayi
Yu, Longxuan
Thakuria, Partha
Brekelmans, Rob
Papalexakis, Evangelos E.
Steeg, Greg Ver
author_facet Wu, Yunshu
Cheng, Jiayi
Yu, Longxuan
Thakuria, Partha
Brekelmans, Rob
Papalexakis, Evangelos E.
Steeg, Greg Ver
contents Continuous diffusion is a natural framework for non-autoregressive generation but has generally lagged behind masked discrete diffusion models (MDMs) on discrete sequence generation. We argue that the bottleneck is not continuity itself, but a representation in which denoising depends on timestep-indexed noise regimes. We introduce \emph{Discrete Stochastic Localization} (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR) under the localization channel. One trained network then supports an entire family of per-token SNR paths, with endpoint masked-diffusion paths as a special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from $T{=}128$ to $T{=}1024$, and the same checkpoint supports random-order autoregressive sampling, as well as a hybrid continuous-then-discrete sampler using as few as T=48 total steps -- without distillation or retraining.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12836
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Discrete Stochastic Localization for Non-autoregressive Generation
Wu, Yunshu
Cheng, Jiayi
Yu, Longxuan
Thakuria, Partha
Brekelmans, Rob
Papalexakis, Evangelos E.
Steeg, Greg Ver
Machine Learning
Continuous diffusion is a natural framework for non-autoregressive generation but has generally lagged behind masked discrete diffusion models (MDMs) on discrete sequence generation. We argue that the bottleneck is not continuity itself, but a representation in which denoising depends on timestep-indexed noise regimes. We introduce \emph{Discrete Stochastic Localization} (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR) under the localization channel. One trained network then supports an entire family of per-token SNR paths, with endpoint masked-diffusion paths as a special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from $T{=}128$ to $T{=}1024$, and the same checkpoint supports random-order autoregressive sampling, as well as a hybrid continuous-then-discrete sampler using as few as T=48 total steps -- without distillation or retraining.
title Discrete Stochastic Localization for Non-autoregressive Generation
topic Machine Learning
url https://arxiv.org/abs/2605.12836