DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910349894090752 |
|---|---|
| author | Huybrechts, Goeric Ronanki, Srikanth Li, Xilai Nosrati, Hadis Bodapati, Sravan Kirchhoff, Katrin |
| author_facet | Huybrechts, Goeric Ronanki, Srikanth Li, Xilai Nosrati, Hadis Bodapati, Sravan Kirchhoff, Katrin |
| contents | Conformer-based end-to-end models have become ubiquitous these days and are commonly used in both streaming and non-streaming automatic speech recognition (ASR). Techniques like dual-mode and dynamic chunk training helped unify streaming and non-streaming systems. However, there remains a performance gap between streaming with a full and limited past context. To address this issue, we propose the integration of a novel dynamic contextual carry-over mechanism in a state-of-the-art (SOTA) unified ASR system. Our proposed dynamic context Conformer (DCTX-Conformer) utilizes a non-overlapping contextual carry-over mechanism that takes into account both the left context of a chunk and one or more preceding context embeddings. We outperform the SOTA by a relative 25.0% word error rate, with a negligible latency impact due to the additional context embeddings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2306_08175 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR Huybrechts, Goeric Ronanki, Srikanth Li, Xilai Nosrati, Hadis Bodapati, Sravan Kirchhoff, Katrin Audio and Speech Processing Artificial Intelligence Machine Learning Sound Conformer-based end-to-end models have become ubiquitous these days and are commonly used in both streaming and non-streaming automatic speech recognition (ASR). Techniques like dual-mode and dynamic chunk training helped unify streaming and non-streaming systems. However, there remains a performance gap between streaming with a full and limited past context. To address this issue, we propose the integration of a novel dynamic contextual carry-over mechanism in a state-of-the-art (SOTA) unified ASR system. Our proposed dynamic context Conformer (DCTX-Conformer) utilizes a non-overlapping contextual carry-over mechanism that takes into account both the left context of a chunk and one or more preceding context embeddings. We outperform the SOTA by a relative 25.0% word error rate, with a negligible latency impact due to the additional context embeddings. |
| title | DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR |
| topic | Audio and Speech Processing Artificial Intelligence Machine Learning Sound |
| url | https://arxiv.org/abs/2306.08175 |