Discrete Stochastic Localization for Non-autoregressive Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914584394203136 |
|---|---|
| author | Wu, Yunshu Cheng, Jiayi Yu, Longxuan Thakuria, Partha Brekelmans, Rob Papalexakis, Evangelos E. Steeg, Greg Ver |
| author_facet | Wu, Yunshu Cheng, Jiayi Yu, Longxuan Thakuria, Partha Brekelmans, Rob Papalexakis, Evangelos E. Steeg, Greg Ver |
| contents | Continuous diffusion is a natural framework for non-autoregressive generation but has generally lagged behind masked discrete diffusion models (MDMs) on discrete sequence generation. We argue that the bottleneck is not continuity itself, but a representation in which denoising depends on timestep-indexed noise regimes. We introduce \emph{Discrete Stochastic Localization} (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR) under the localization channel. One trained network then supports an entire family of per-token SNR paths, with endpoint masked-diffusion paths as a special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from $T{=}128$ to $T{=}1024$, and the same checkpoint supports random-order autoregressive sampling, as well as a hybrid continuous-then-discrete sampler using as few as T=48 total steps -- without distillation or retraining. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_12836 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Discrete Stochastic Localization for Non-autoregressive Generation Wu, Yunshu Cheng, Jiayi Yu, Longxuan Thakuria, Partha Brekelmans, Rob Papalexakis, Evangelos E. Steeg, Greg Ver Machine Learning Continuous diffusion is a natural framework for non-autoregressive generation but has generally lagged behind masked discrete diffusion models (MDMs) on discrete sequence generation. We argue that the bottleneck is not continuity itself, but a representation in which denoising depends on timestep-indexed noise regimes. We introduce \emph{Discrete Stochastic Localization} (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR) under the localization channel. One trained network then supports an entire family of per-token SNR paths, with endpoint masked-diffusion paths as a special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from $T{=}128$ to $T{=}1024$, and the same checkpoint supports random-order autoregressive sampling, as well as a hybrid continuous-then-discrete sampler using as few as T=48 total steps -- without distillation or retraining. |
| title | Discrete Stochastic Localization for Non-autoregressive Generation |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2605.12836 |