SLM-S2ST: A multimodal language model for direct speech-to-speech translation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911438889549824 |
|---|---|
| author | Hu, Yuxuan Wu, Haibin Fan, Ruchao Wang, Xiaofei Lu, Heng Qian, Yao Li, Jinyu |
| author_facet | Hu, Yuxuan Wu, Haibin Fan, Ruchao Wang, Xiaofei Lu, Heng Qian, Yao Li, Jinyu |
| contents | Speech-aware language models (LMs) have demonstrated capabilities in understanding spoken language while generating text-based responses. However, enabling them to produce speech output efficiently and effectively remains a challenge. In this paper, we present SLM-S2ST, a multimodal LM for direct speech-to-speech translation (S2ST), built on the open-source Phi4-MM model. SLM-S2ST extends its predecessor by generating translated speech using an audio transformer head that predicts audio tokens with a delay relative to text tokens, followed by a streaming vocoder for waveform synthesis. Our experimental results on the CVSS-C dataset demonstrate SLM-S2ST's superior performance, significantly surpassing existing baseline models trained on the same dataset. Furthermore, when we scale up the training data and the model size, SLM-S2ST reaches on-par performance with the current SOTA model. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_04392 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SLM-S2ST: A multimodal language model for direct speech-to-speech translation Hu, Yuxuan Wu, Haibin Fan, Ruchao Wang, Xiaofei Lu, Heng Qian, Yao Li, Jinyu Audio and Speech Processing Speech-aware language models (LMs) have demonstrated capabilities in understanding spoken language while generating text-based responses. However, enabling them to produce speech output efficiently and effectively remains a challenge. In this paper, we present SLM-S2ST, a multimodal LM for direct speech-to-speech translation (S2ST), built on the open-source Phi4-MM model. SLM-S2ST extends its predecessor by generating translated speech using an audio transformer head that predicts audio tokens with a delay relative to text tokens, followed by a streaming vocoder for waveform synthesis. Our experimental results on the CVSS-C dataset demonstrate SLM-S2ST's superior performance, significantly surpassing existing baseline models trained on the same dataset. Furthermore, when we scale up the training data and the model size, SLM-S2ST reaches on-par performance with the current SOTA model. |
| title | SLM-S2ST: A multimodal language model for direct speech-to-speech translation |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.04392 |