SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Papi, Sara, Gaido, Marco, Negri, Matteo, Bentivogli, Luisa
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911927407476736
author Papi, Sara
Gaido, Marco
Negri, Matteo
Bentivogli, Luisa
author_facet Papi, Sara
Gaido, Marco
Negri, Matteo
Bentivogli, Luisa
contents This paper describes the FBK's participation in the Simultaneous Translation Evaluation Campaign at IWSLT 2024. For this year's submission in the speech-to-text translation (ST) sub-track, we propose SimulSeamless, which is realized by combining AlignAtt and SeamlessM4T in its medium configuration. The SeamlessM4T model is used "off-the-shelf" and its simultaneous inference is enabled through the adoption of AlignAtt, a SimulST policy based on cross-attention that can be applied without any retraining or adaptation of the underlying model for the simultaneous task. We participated in all the Shared Task languages (English->{German, Japanese, Chinese}, and Czech->English), achieving acceptable or even better results compared to last year's submissions. SimulSeamless, covering more than 143 source languages and 200 target languages, is released at: https://github.com/hlt-mt/FBK-fairseq/.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14177
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
Papi, Sara
Gaido, Marco
Negri, Matteo
Bentivogli, Luisa
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
This paper describes the FBK's participation in the Simultaneous Translation Evaluation Campaign at IWSLT 2024. For this year's submission in the speech-to-text translation (ST) sub-track, we propose SimulSeamless, which is realized by combining AlignAtt and SeamlessM4T in its medium configuration. The SeamlessM4T model is used "off-the-shelf" and its simultaneous inference is enabled through the adoption of AlignAtt, a SimulST policy based on cross-attention that can be applied without any retraining or adaptation of the underlying model for the simultaneous task. We participated in all the Shared Task languages (English->{German, Japanese, Chinese}, and Czech->English), achieving acceptable or even better results compared to last year's submissions. SimulSeamless, covering more than 143 source languages and 200 target languages, is released at: https://github.com/hlt-mt/FBK-fairseq/.
title SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.14177