Streaming Speech-to-Confusion Network Speech Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Filimonov, Denis, Pandey, Prabhat, Rastrow, Ariya, Gandhe, Ankur, Stolcke, Andreas
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910307974119424
author Filimonov, Denis
Pandey, Prabhat
Rastrow, Ariya
Gandhe, Ankur
Stolcke, Andreas
author_facet Filimonov, Denis
Pandey, Prabhat
Rastrow, Ariya
Gandhe, Ankur
Stolcke, Andreas
contents In interactive automatic speech recognition (ASR) systems, low-latency requirements limit the amount of search space that can be explored during decoding, particularly in end-to-end neural ASR. In this paper, we present a novel streaming ASR architecture that outputs a confusion network while maintaining limited latency, as needed for interactive applications. We show that 1-best results of our model are on par with a comparable RNN-T system, while the richer hypothesis set allows second-pass rescoring to achieve 10-20\% lower word error rate on the LibriSpeech task. We also show that our model outperforms a strong RNN-T baseline on a far-field voice assistant task.
format Preprint
id arxiv_https___arxiv_org_abs_2306_03778
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Streaming Speech-to-Confusion Network Speech Recognition
Filimonov, Denis
Pandey, Prabhat
Rastrow, Ariya
Gandhe, Ankur
Stolcke, Andreas
Audio and Speech Processing
Computation and Language
In interactive automatic speech recognition (ASR) systems, low-latency requirements limit the amount of search space that can be explored during decoding, particularly in end-to-end neural ASR. In this paper, we present a novel streaming ASR architecture that outputs a confusion network while maintaining limited latency, as needed for interactive applications. We show that 1-best results of our model are on par with a comparable RNN-T system, while the richer hypothesis set allows second-pass rescoring to achieve 10-20\% lower word error rate on the LibriSpeech task. We also show that our model outperforms a strong RNN-T baseline on a far-field voice assistant task.
title Streaming Speech-to-Confusion Network Speech Recognition
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2306.03778