CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Rui, Li, Jinyu, Fan, Ruchao, Post, Matt
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913535645188096
author Zhao, Rui
Li, Jinyu
Fan, Ruchao
Post, Matt
author_facet Zhao, Rui
Li, Jinyu
Fan, Ruchao
Post, Matt
contents Models for streaming speech translation (ST) can achieve high accuracy and low latency if they're developed with vast amounts of paired audio in the source language and written text in the target language. Yet, these text labels for the target language are often pseudo labels due to the prohibitive cost of manual ST data labeling. In this paper, we introduce a methodology named Connectionist Temporal Classification guided modality matching (CTC-GMM) that enhances the streaming ST model by leveraging extensive machine translation (MT) text data. This technique employs CTC to compress the speech sequence into a compact embedding sequence that matches the corresponding text sequence, allowing us to utilize matched {source-target} language text pairs from the MT corpora to refine the streaming ST model further. Our evaluations with FLEURS and CoVoST2 show that the CTC-GMM approach can increase translation accuracy relatively by 13.9% and 6.4% respectively, while also boosting decoding speed by 59.7% on GPU.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05146
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
Zhao, Rui
Li, Jinyu
Fan, Ruchao
Post, Matt
Computation and Language
Artificial Intelligence
Audio and Speech Processing
Models for streaming speech translation (ST) can achieve high accuracy and low latency if they're developed with vast amounts of paired audio in the source language and written text in the target language. Yet, these text labels for the target language are often pseudo labels due to the prohibitive cost of manual ST data labeling. In this paper, we introduce a methodology named Connectionist Temporal Classification guided modality matching (CTC-GMM) that enhances the streaming ST model by leveraging extensive machine translation (MT) text data. This technique employs CTC to compress the speech sequence into a compact embedding sequence that matches the corresponding text sequence, allowing us to utilize matched {source-target} language text pairs from the MT corpora to refine the streaming ST model further. Our evaluations with FLEURS and CoVoST2 show that the CTC-GMM approach can increase translation accuracy relatively by 13.9% and 6.4% respectively, while also boosting decoding speed by 59.7% on GPU.
title CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
topic Computation and Language
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2410.05146