Text-Utilization for Encoder-dominated Speech Recognition Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zeyer, Albert, Posielek, Tim, Schlüter, Ralf, Ney, Hermann
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915967489015808
author Zeyer, Albert
Posielek, Tim
Schlüter, Ralf
Ney, Hermann
author_facet Zeyer, Albert
Posielek, Tim
Schlüter, Ralf
Ney, Hermann
contents This paper investigates efficient methods for utilizing text-only data to improve speech recognition, focusing on encoder-dominated models that facilitate faster recognition. We provide a comprehensive comparison of techniques to integrate text-only data, including modality matching and dynamic downsampling to reach text-level representations within the encoder. Our experiments on the LibriSpeech corpus show that a larger encoder with a smaller decoder can equal or surpass the performance of architectures with larger decoders. We demonstrate that simple configurations, such as random duration models, are often more effective than complex alternatives, significantly simplifying the training pipeline. All code and recipes are made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26514
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Text-Utilization for Encoder-dominated Speech Recognition Models
Zeyer, Albert
Posielek, Tim
Schlüter, Ralf
Ney, Hermann
Computation and Language
Artificial Intelligence
Neural and Evolutionary Computing
This paper investigates efficient methods for utilizing text-only data to improve speech recognition, focusing on encoder-dominated models that facilitate faster recognition. We provide a comprehensive comparison of techniques to integrate text-only data, including modality matching and dynamic downsampling to reach text-level representations within the encoder. Our experiments on the LibriSpeech corpus show that a larger encoder with a smaller decoder can equal or surpass the performance of architectures with larger decoders. We demonstrate that simple configurations, such as random duration models, are often more effective than complex alternatives, significantly simplifying the training pipeline. All code and recipes are made publicly available.
title Text-Utilization for Encoder-dominated Speech Recognition Models
topic Computation and Language
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2604.26514