Task Arithmetic for Language Expansion in Speech Translation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Yao-Fei, Futami, Hayato, Kashiwagi, Yosuke, Tsunoo, Emiru, Teo, Wen Shen, Arora, Siddhant, Watanabe, Shinji
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916867556245504
author Cheng, Yao-Fei
Futami, Hayato
Kashiwagi, Yosuke
Tsunoo, Emiru
Teo, Wen Shen
Arora, Siddhant
Watanabe, Shinji
author_facet Cheng, Yao-Fei
Futami, Hayato
Kashiwagi, Yosuke
Tsunoo, Emiru
Teo, Wen Shen
Arora, Siddhant
Watanabe, Shinji
contents Recent progress in large language models (LLMs) has gained interest in speech-text multimodal foundation models, achieving strong performance on instruction-tuned speech translation (ST). However, expanding language pairs is costly due to re-training on combined new and previous datasets. To address this, we aim to build a one-to-many ST system from existing one-to-one ST systems using task arithmetic without re-training. Direct application of task arithmetic in ST leads to language confusion; therefore, we introduce an augmented task arithmetic method incorporating a language control model to ensure correct target language generation. Our experiments on MuST-C and CoVoST-2 show BLEU score improvements of up to 4.66 and 4.92, with COMET gains of 8.87 and 11.83. In addition, we demonstrate our framework can extend to language pairs lacking paired ST training data or pre-trained ST models by synthesizing ST models based on existing machine translation (MT) and ST models via task analogies.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11274
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Task Arithmetic for Language Expansion in Speech Translation
Cheng, Yao-Fei
Futami, Hayato
Kashiwagi, Yosuke
Tsunoo, Emiru
Teo, Wen Shen
Arora, Siddhant
Watanabe, Shinji
Computation and Language
Artificial Intelligence
Recent progress in large language models (LLMs) has gained interest in speech-text multimodal foundation models, achieving strong performance on instruction-tuned speech translation (ST). However, expanding language pairs is costly due to re-training on combined new and previous datasets. To address this, we aim to build a one-to-many ST system from existing one-to-one ST systems using task arithmetic without re-training. Direct application of task arithmetic in ST leads to language confusion; therefore, we introduce an augmented task arithmetic method incorporating a language control model to ensure correct target language generation. Our experiments on MuST-C and CoVoST-2 show BLEU score improvements of up to 4.66 and 4.92, with COMET gains of 8.87 and 11.83. In addition, we demonstrate our framework can extend to language pairs lacking paired ST training data or pre-trained ST models by synthesizing ST models based on existing machine translation (MT) and ST models via task analogies.
title Task Arithmetic for Language Expansion in Speech Translation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.11274