Textless Speech-to-Speech Translation With Limited Parallel Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Diwan, Anuj, Srinivasan, Anirudh, Harwath, David, Choi, Eunsol
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915007915098112
author Diwan, Anuj
Srinivasan, Anirudh
Harwath, David
Choi, Eunsol
author_facet Diwan, Anuj
Srinivasan, Anirudh
Harwath, David
Choi, Eunsol
contents Existing speech-to-speech translation (S2ST) models fall into two camps: they either leverage text as an intermediate step or require hundreds of hours of parallel speech data. Both approaches are incompatible with textless languages or language pairs with limited parallel data. We present PFB, a framework for training textless S2ST models that require just dozens of hours of parallel speech data. We first pretrain a model on large-scale monolingual speech data, finetune it with a small amount of parallel speech data (20-60 hours), and lastly train with an unsupervised backtranslation objective. We train and evaluate our models for English-to-German, German-to-English and Marathi-to-English translation on three different domains (European Parliament, Common Voice, and All India Radio) with single-speaker synthesized speech. Evaluated using the ASR-BLEU metric, our models achieve reasonable performance on all three domains, with some being within 1-2 points of our higher-resourced topline.
format Preprint
id arxiv_https___arxiv_org_abs_2305_15405
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Textless Speech-to-Speech Translation With Limited Parallel Data
Diwan, Anuj
Srinivasan, Anirudh
Harwath, David
Choi, Eunsol
Computation and Language
Audio and Speech Processing
Existing speech-to-speech translation (S2ST) models fall into two camps: they either leverage text as an intermediate step or require hundreds of hours of parallel speech data. Both approaches are incompatible with textless languages or language pairs with limited parallel data. We present PFB, a framework for training textless S2ST models that require just dozens of hours of parallel speech data. We first pretrain a model on large-scale monolingual speech data, finetune it with a small amount of parallel speech data (20-60 hours), and lastly train with an unsupervised backtranslation objective. We train and evaluate our models for English-to-German, German-to-English and Marathi-to-English translation on three different domains (European Parliament, Common Voice, and All India Radio) with single-speaker synthesized speech. Evaluated using the ASR-BLEU metric, our models achieve reasonable performance on all three domains, with some being within 1-2 points of our higher-resourced topline.
title Textless Speech-to-Speech Translation With Limited Parallel Data
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2305.15405