Streaming Speech-to-Confusion Network Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910307974119424 |
|---|---|
| author | Filimonov, Denis Pandey, Prabhat Rastrow, Ariya Gandhe, Ankur Stolcke, Andreas |
| author_facet | Filimonov, Denis Pandey, Prabhat Rastrow, Ariya Gandhe, Ankur Stolcke, Andreas |
| contents | In interactive automatic speech recognition (ASR) systems, low-latency requirements limit the amount of search space that can be explored during decoding, particularly in end-to-end neural ASR. In this paper, we present a novel streaming ASR architecture that outputs a confusion network while maintaining limited latency, as needed for interactive applications. We show that 1-best results of our model are on par with a comparable RNN-T system, while the richer hypothesis set allows second-pass rescoring to achieve 10-20\% lower word error rate on the LibriSpeech task. We also show that our model outperforms a strong RNN-T baseline on a far-field voice assistant task. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2306_03778 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Streaming Speech-to-Confusion Network Speech Recognition Filimonov, Denis Pandey, Prabhat Rastrow, Ariya Gandhe, Ankur Stolcke, Andreas Audio and Speech Processing Computation and Language In interactive automatic speech recognition (ASR) systems, low-latency requirements limit the amount of search space that can be explored during decoding, particularly in end-to-end neural ASR. In this paper, we present a novel streaming ASR architecture that outputs a confusion network while maintaining limited latency, as needed for interactive applications. We show that 1-best results of our model are on par with a comparable RNN-T system, while the richer hypothesis set allows second-pass rescoring to achieve 10-20\% lower word error rate on the LibriSpeech task. We also show that our model outperforms a strong RNN-T baseline on a far-field voice assistant task. |
| title | Streaming Speech-to-Confusion Network Speech Recognition |
| topic | Audio and Speech Processing Computation and Language |
| url | https://arxiv.org/abs/2306.03778 |