POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xuanchen, Cui, Chenrui, Wang, Tianrui, Ge, Meng, Huang, Zikang, Peng, Yizhou, Li, Jin, Lu, Yuheng, Jiang, Yu, Tashi, Nyima, Wang, Longbiao, Dang, Jianwu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908922693025792
author Li, Xuanchen
Cui, Chenrui
Wang, Tianrui
Ge, Meng
Huang, Zikang
Peng, Yizhou
Li, Jin
Lu, Yuheng
Jiang, Yu
Tashi, Nyima
Wang, Longbiao
Dang, Jianwu
author_facet Li, Xuanchen
Cui, Chenrui
Wang, Tianrui
Ge, Meng
Huang, Zikang
Peng, Yizhou
Li, Jin
Lu, Yuheng
Jiang, Yu
Tashi, Nyima
Wang, Longbiao
Dang, Jianwu
contents Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation. However, existing approaches often overlook semantic commonalities across source languages, leading to biased translation performance. In this work, we propose POTSA (Parallel Optimal Transport for Speech Alignment), a new framework based on cross-lingual parallel speech pairs and Optimal Transport, designed to bridge high- and low-resource translation gaps. First, we introduce a Bias Compensation module to coarsely align initial speech representations. Second, we impose token-level OT constraints on a Q-Former using parallel pairs to establish fine-grained representation consistency. Then, we apply a layer scheduling strategy to focus OT constraints on semantically beneficial layers. Experiments on FLEURS show our method achieves SOTA performance, with +1.29 BLEU over five common languages and +2.93 BLEU on zero-shot languages, using only 10 hours of parallel speech per language.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09232
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
Li, Xuanchen
Cui, Chenrui
Wang, Tianrui
Ge, Meng
Huang, Zikang
Peng, Yizhou
Li, Jin
Lu, Yuheng
Jiang, Yu
Tashi, Nyima
Wang, Longbiao
Dang, Jianwu
Computation and Language
Sound
Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation. However, existing approaches often overlook semantic commonalities across source languages, leading to biased translation performance. In this work, we propose POTSA (Parallel Optimal Transport for Speech Alignment), a new framework based on cross-lingual parallel speech pairs and Optimal Transport, designed to bridge high- and low-resource translation gaps. First, we introduce a Bias Compensation module to coarsely align initial speech representations. Second, we impose token-level OT constraints on a Q-Former using parallel pairs to establish fine-grained representation consistency. Then, we apply a layer scheduling strategy to focus OT constraints on semantically beneficial layers. Experiments on FLEURS show our method achieves SOTA performance, with +1.29 BLEU over five common languages and +2.93 BLEU on zero-shot languages, using only 10 hours of parallel speech per language.
title POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
topic Computation and Language
Sound
url https://arxiv.org/abs/2511.09232