Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Moslem, Yasmin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929401190416384
author Moslem, Yasmin
author_facet Moslem, Yasmin
contents This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2024) for Irish-to-English speech translation. We built end-to-end systems based on Whisper, and employed a number of data augmentation techniques, such as speech back-translation and noise augmentation. We investigate the effect of using synthetic audio data and discuss several methods for enriching signal diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17363
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
Moslem, Yasmin
Computation and Language
Sound
Audio and Speech Processing
This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2024) for Irish-to-English speech translation. We built end-to-end systems based on Whisper, and employed a number of data augmentation techniques, such as speech back-translation and noise augmentation. We investigate the effect of using synthetic audio data and discuss several methods for enriching signal diversity.
title Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.17363