SLM-S2ST: A multimodal language model for direct speech-to-speech translation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Yuxuan, Wu, Haibin, Fan, Ruchao, Wang, Xiaofei, Lu, Heng, Qian, Yao, Li, Jinyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911438889549824
author Hu, Yuxuan
Wu, Haibin
Fan, Ruchao
Wang, Xiaofei
Lu, Heng
Qian, Yao
Li, Jinyu
author_facet Hu, Yuxuan
Wu, Haibin
Fan, Ruchao
Wang, Xiaofei
Lu, Heng
Qian, Yao
Li, Jinyu
contents Speech-aware language models (LMs) have demonstrated capabilities in understanding spoken language while generating text-based responses. However, enabling them to produce speech output efficiently and effectively remains a challenge. In this paper, we present SLM-S2ST, a multimodal LM for direct speech-to-speech translation (S2ST), built on the open-source Phi4-MM model. SLM-S2ST extends its predecessor by generating translated speech using an audio transformer head that predicts audio tokens with a delay relative to text tokens, followed by a streaming vocoder for waveform synthesis. Our experimental results on the CVSS-C dataset demonstrate SLM-S2ST's superior performance, significantly surpassing existing baseline models trained on the same dataset. Furthermore, when we scale up the training data and the model size, SLM-S2ST reaches on-par performance with the current SOTA model.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04392
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SLM-S2ST: A multimodal language model for direct speech-to-speech translation
Hu, Yuxuan
Wu, Haibin
Fan, Ruchao
Wang, Xiaofei
Lu, Heng
Qian, Yao
Li, Jinyu
Audio and Speech Processing
Speech-aware language models (LMs) have demonstrated capabilities in understanding spoken language while generating text-based responses. However, enabling them to produce speech output efficiently and effectively remains a challenge. In this paper, we present SLM-S2ST, a multimodal LM for direct speech-to-speech translation (S2ST), built on the open-source Phi4-MM model. SLM-S2ST extends its predecessor by generating translated speech using an audio transformer head that predicts audio tokens with a delay relative to text tokens, followed by a streaming vocoder for waveform synthesis. Our experimental results on the CVSS-C dataset demonstrate SLM-S2ST's superior performance, significantly surpassing existing baseline models trained on the same dataset. Furthermore, when we scale up the training data and the model size, SLM-S2ST reaches on-par performance with the current SOTA model.
title SLM-S2ST: A multimodal language model for direct speech-to-speech translation
topic Audio and Speech Processing
url https://arxiv.org/abs/2506.04392