Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Futami, Hayato, Tsunoo, Emiru, Kashiwagi, Yosuke, Ito, Yuki, Shahmohammadi, Hassan, Arora, Siddhant, Watanabe, Shinji
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916791762026496
author Futami, Hayato
Tsunoo, Emiru
Kashiwagi, Yosuke
Ito, Yuki
Shahmohammadi, Hassan
Arora, Siddhant
Watanabe, Shinji
author_facet Futami, Hayato
Tsunoo, Emiru
Kashiwagi, Yosuke
Ito, Yuki
Shahmohammadi, Hassan
Arora, Siddhant
Watanabe, Shinji
contents Speech-to-speech translation (S2ST) has been advanced with large language models (LLMs), which are fine-tuned on discrete speech units. In such approaches, modality adaptation from text to speech has been an issue. LLMs are trained on text-only data, which presents challenges to adapt them to speech modality with limited speech-to-speech data. To address the training difficulty, we propose scheduled interleaved speech--text training in this study. We use interleaved speech--text units instead of speech units during training, where aligned text tokens are interleaved at the word level. We gradually decrease the ratio of text as training progresses, to facilitate progressive modality adaptation from text to speech. We conduct experimental evaluations by fine-tuning LLaMA3.2-1B for S2ST on the CVSS dataset. We show that the proposed method consistently improves the translation performances, especially for languages with limited training data.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
Futami, Hayato
Tsunoo, Emiru
Kashiwagi, Yosuke
Ito, Yuki
Shahmohammadi, Hassan
Arora, Siddhant
Watanabe, Shinji
Computation and Language
Sound
Audio and Speech Processing
Speech-to-speech translation (S2ST) has been advanced with large language models (LLMs), which are fine-tuned on discrete speech units. In such approaches, modality adaptation from text to speech has been an issue. LLMs are trained on text-only data, which presents challenges to adapt them to speech modality with limited speech-to-speech data. To address the training difficulty, we propose scheduled interleaved speech--text training in this study. We use interleaved speech--text units instead of speech units during training, where aligned text tokens are interleaved at the word level. We gradually decrease the ratio of text as training progresses, to facilitate progressive modality adaptation from text to speech. We conduct experimental evaluations by fine-tuning LLaMA3.2-1B for S2ST on the CVSS dataset. We show that the proposed method consistently improves the translation performances, especially for languages with limited training data.
title Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.10299