Diffusion Language Models for Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915966916493312 |
|---|---|
| author | Naveriani, Davyd Zeyer, Albert Schlüter, Ralf Ney, Hermann |
| author_facet | Naveriani, Davyd Zeyer, Albert Schlüter, Ralf Ney, Hermann |
| contents | Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporating masked diffusion language models (MDLM) and uniform-state diffusion models (USDMs) for rescoring ASR hypotheses. Additionally, we design a new joint-decoding method that combines CTC and USDM by integrating the framewise probability distributions derived from CTC with the labelwise probability distributions computed by USDM at each decoding step, thereby generating new candidates that combine strong language knowledge from USDM and acoustic information from CTC. Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text. We publish all our code and recipes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_14001 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Diffusion Language Models for Speech Recognition Naveriani, Davyd Zeyer, Albert Schlüter, Ralf Ney, Hermann Computation and Language Artificial Intelligence Machine Learning Neural and Evolutionary Computing Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporating masked diffusion language models (MDLM) and uniform-state diffusion models (USDMs) for rescoring ASR hypotheses. Additionally, we design a new joint-decoding method that combines CTC and USDM by integrating the framewise probability distributions derived from CTC with the labelwise probability distributions computed by USDM at each decoding step, thereby generating new candidates that combine strong language knowledge from USDM and acoustic information from CTC. Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text. We publish all our code and recipes. |
| title | Diffusion Language Models for Speech Recognition |
| topic | Computation and Language Artificial Intelligence Machine Learning Neural and Evolutionary Computing |
| url | https://arxiv.org/abs/2604.14001 |