Diffusion Language Models for Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Naveriani, Davyd, Zeyer, Albert, Schlüter, Ralf, Ney, Hermann
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915966916493312
author Naveriani, Davyd
Zeyer, Albert
Schlüter, Ralf
Ney, Hermann
author_facet Naveriani, Davyd
Zeyer, Albert
Schlüter, Ralf
Ney, Hermann
contents Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporating masked diffusion language models (MDLM) and uniform-state diffusion models (USDMs) for rescoring ASR hypotheses. Additionally, we design a new joint-decoding method that combines CTC and USDM by integrating the framewise probability distributions derived from CTC with the labelwise probability distributions computed by USDM at each decoding step, thereby generating new candidates that combine strong language knowledge from USDM and acoustic information from CTC. Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text. We publish all our code and recipes.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14001
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Diffusion Language Models for Speech Recognition
Naveriani, Davyd
Zeyer, Albert
Schlüter, Ralf
Ney, Hermann
Computation and Language
Artificial Intelligence
Machine Learning
Neural and Evolutionary Computing
Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporating masked diffusion language models (MDLM) and uniform-state diffusion models (USDMs) for rescoring ASR hypotheses. Additionally, we design a new joint-decoding method that combines CTC and USDM by integrating the framewise probability distributions derived from CTC with the labelwise probability distributions computed by USDM at each decoding step, thereby generating new candidates that combine strong language knowledge from USDM and acoustic information from CTC. Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text. We publish all our code and recipes.
title Diffusion Language Models for Speech Recognition
topic Computation and Language
Artificial Intelligence
Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2604.14001