DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huybrechts, Goeric, Ronanki, Srikanth, Li, Xilai, Nosrati, Hadis, Bodapati, Sravan, Kirchhoff, Katrin
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910349894090752
author Huybrechts, Goeric
Ronanki, Srikanth
Li, Xilai
Nosrati, Hadis
Bodapati, Sravan
Kirchhoff, Katrin
author_facet Huybrechts, Goeric
Ronanki, Srikanth
Li, Xilai
Nosrati, Hadis
Bodapati, Sravan
Kirchhoff, Katrin
contents Conformer-based end-to-end models have become ubiquitous these days and are commonly used in both streaming and non-streaming automatic speech recognition (ASR). Techniques like dual-mode and dynamic chunk training helped unify streaming and non-streaming systems. However, there remains a performance gap between streaming with a full and limited past context. To address this issue, we propose the integration of a novel dynamic contextual carry-over mechanism in a state-of-the-art (SOTA) unified ASR system. Our proposed dynamic context Conformer (DCTX-Conformer) utilizes a non-overlapping contextual carry-over mechanism that takes into account both the left context of a chunk and one or more preceding context embeddings. We outperform the SOTA by a relative 25.0% word error rate, with a negligible latency impact due to the additional context embeddings.
format Preprint
id arxiv_https___arxiv_org_abs_2306_08175
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
Huybrechts, Goeric
Ronanki, Srikanth
Li, Xilai
Nosrati, Hadis
Bodapati, Sravan
Kirchhoff, Katrin
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
Conformer-based end-to-end models have become ubiquitous these days and are commonly used in both streaming and non-streaming automatic speech recognition (ASR). Techniques like dual-mode and dynamic chunk training helped unify streaming and non-streaming systems. However, there remains a performance gap between streaming with a full and limited past context. To address this issue, we propose the integration of a novel dynamic contextual carry-over mechanism in a state-of-the-art (SOTA) unified ASR system. Our proposed dynamic context Conformer (DCTX-Conformer) utilizes a non-overlapping contextual carry-over mechanism that takes into account both the left context of a chunk and one or more preceding context embeddings. We outperform the SOTA by a relative 25.0% word error rate, with a negligible latency impact due to the additional context embeddings.
title DCTX-Conformer: Dynamic context carry-over for low latency unified streaming and non-streaming Conformer ASR
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
url https://arxiv.org/abs/2306.08175