Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Yuhao, Ma, Xiangnan, Kou, Kaiqi, Liu, Peizhuo, Shan, Weiqiao, Wang, Benyou, Xiao, Tong, Huang, Yuxin, Yu, Zhengtao, Zhu, Jingbo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918028870942720
author Zhang, Yuhao
Ma, Xiangnan
Kou, Kaiqi
Liu, Peizhuo
Shan, Weiqiao
Wang, Benyou
Xiao, Tong
Huang, Yuxin
Yu, Zhengtao
Zhu, Jingbo
author_facet Zhang, Yuhao
Ma, Xiangnan
Kou, Kaiqi
Liu, Peizhuo
Shan, Weiqiao
Wang, Benyou
Xiao, Tong
Huang, Yuxin
Yu, Zhengtao
Zhu, Jingbo
contents The success of building textless speech-to-speech translation (S2ST) models has attracted much attention. However, S2ST still faces two main challenges: 1) extracting linguistic features for various speech signals, called cross-modal (CM), and 2) learning alignment of difference languages in long sequences, called cross-lingual (CL). We propose the unit language to overcome the two modeling challenges. The unit language can be considered a text-like representation format, constructed using $n$-gram language modeling. We implement multi-task learning to utilize the unit language in guiding the speech modeling process. Our initial results reveal a conflict when applying source and target unit languages simultaneously. We propose task prompt modeling to mitigate this conflict. We conduct experiments on four languages of the Voxpupil dataset. Our method demonstrates significant improvements over a strong baseline and achieves performance comparable to models trained with text.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
Zhang, Yuhao
Ma, Xiangnan
Kou, Kaiqi
Liu, Peizhuo
Shan, Weiqiao
Wang, Benyou
Xiao, Tong
Huang, Yuxin
Yu, Zhengtao
Zhu, Jingbo
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
The success of building textless speech-to-speech translation (S2ST) models has attracted much attention. However, S2ST still faces two main challenges: 1) extracting linguistic features for various speech signals, called cross-modal (CM), and 2) learning alignment of difference languages in long sequences, called cross-lingual (CL). We propose the unit language to overcome the two modeling challenges. The unit language can be considered a text-like representation format, constructed using $n$-gram language modeling. We implement multi-task learning to utilize the unit language in guiding the speech modeling process. Our initial results reveal a conflict when applying source and target unit languages simultaneously. We propose task prompt modeling to mitigate this conflict. We conduct experiments on four languages of the Voxpupil dataset. Our method demonstrates significant improvements over a strong baseline and achieves performance comparable to models trained with text.
title Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.15333