Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Van Gysel, Christophe, Wu, Maggie, Verwimp, Lyan, Tirkaz, Caglar, Bertola, Marco, Lei, Zhihong, Oualil, Youssef
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912417374535680
author Van Gysel, Christophe
Wu, Maggie
Verwimp, Lyan
Tirkaz, Caglar
Bertola, Marco
Lei, Zhihong
Oualil, Youssef
author_facet Van Gysel, Christophe
Wu, Maggie
Verwimp, Lyan
Tirkaz, Caglar
Bertola, Marco
Lei, Zhihong
Oualil, Youssef
contents End-to-end (E2E) Automatic Speech Recognition (ASR) models are trained using paired audio-text samples that are expensive to obtain, since high-quality ground-truth data requires human annotators. Voice search applications, such as digital media players, leverage ASR to allow users to search by voice as opposed to an on-screen keyboard. However, recent or infrequent movie titles may not be sufficiently represented in the E2E ASR system's training data, and hence, may suffer poor recognition. In this paper, we propose a phonetic correction system that consists of (a) a phonetic search based on the ASR model's output that generates phonetic alternatives that may not be considered by the E2E system, and (b) a rescorer component that combines the ASR model recognition and the phonetic alternatives, and select a final system output. We find that our approach improves word error rate between 4.4 and 7.6% relative on benchmarks of popular movie titles over a series of competitive baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06117
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction
Van Gysel, Christophe
Wu, Maggie
Verwimp, Lyan
Tirkaz, Caglar
Bertola, Marco
Lei, Zhihong
Oualil, Youssef
Computation and Language
Artificial Intelligence
Information Retrieval
End-to-end (E2E) Automatic Speech Recognition (ASR) models are trained using paired audio-text samples that are expensive to obtain, since high-quality ground-truth data requires human annotators. Voice search applications, such as digital media players, leverage ASR to allow users to search by voice as opposed to an on-screen keyboard. However, recent or infrequent movie titles may not be sufficiently represented in the E2E ASR system's training data, and hence, may suffer poor recognition. In this paper, we propose a phonetic correction system that consists of (a) a phonetic search based on the ASR model's output that generates phonetic alternatives that may not be considered by the E2E system, and (b) a rescorer component that combines the ASR model recognition and the phonetic alternatives, and select a final system output. We find that our approach improves word error rate between 4.4 and 7.6% relative on benchmarks of popular movie titles over a series of competitive baselines.
title Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2506.06117