Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Peidong, Xue, Jian, Li, Jinyu, Chen, Junkun, Subramanian, Aswin Shanmugam
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909224568619008
author Wang, Peidong
Xue, Jian
Li, Jinyu
Chen, Junkun
Subramanian, Aswin Shanmugam
author_facet Wang, Peidong
Xue, Jian
Li, Jinyu
Chen, Junkun
Subramanian, Aswin Shanmugam
contents Language-agnostic many-to-one end-to-end speech translation models can convert audio signals from different source languages into text in a target language. These models do not need source language identification, which improves user experience. In some cases, the input language can be given or estimated. Our goal is to use this additional language information while preserving the quality of the other languages. We accomplish this by introducing a simple and effective linear input network. The linear input network is initialized as an identity matrix, which ensures that the model can perform as well as, or better than, the original model. Experimental results show that the proposed method can successfully enhance the specified language, while keeping the language-agnostic ability of the many-to-one ST models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10276
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
Wang, Peidong
Xue, Jian
Li, Jinyu
Chen, Junkun
Subramanian, Aswin Shanmugam
Computation and Language
Sound
Audio and Speech Processing
Language-agnostic many-to-one end-to-end speech translation models can convert audio signals from different source languages into text in a target language. These models do not need source language identification, which improves user experience. In some cases, the input language can be given or estimated. Our goal is to use this additional language information while preserving the quality of the other languages. We accomplish this by introducing a simple and effective linear input network. The linear input network is initialized as an identity matrix, which ensures that the model can perform as well as, or better than, the original model. Experimental results show that the proposed method can successfully enhance the specified language, while keeping the language-agnostic ability of the many-to-one ST models.
title Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.10276