POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908922693025792 |
|---|---|
| author | Li, Xuanchen Cui, Chenrui Wang, Tianrui Ge, Meng Huang, Zikang Peng, Yizhou Li, Jin Lu, Yuheng Jiang, Yu Tashi, Nyima Wang, Longbiao Dang, Jianwu |
| author_facet | Li, Xuanchen Cui, Chenrui Wang, Tianrui Ge, Meng Huang, Zikang Peng, Yizhou Li, Jin Lu, Yuheng Jiang, Yu Tashi, Nyima Wang, Longbiao Dang, Jianwu |
| contents | Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation. However, existing approaches often overlook semantic commonalities across source languages, leading to biased translation performance. In this work, we propose POTSA (Parallel Optimal Transport for Speech Alignment), a new framework based on cross-lingual parallel speech pairs and Optimal Transport, designed to bridge high- and low-resource translation gaps. First, we introduce a Bias Compensation module to coarsely align initial speech representations. Second, we impose token-level OT constraints on a Q-Former using parallel pairs to establish fine-grained representation consistency. Then, we apply a layer scheduling strategy to focus OT constraints on semantically beneficial layers. Experiments on FLEURS show our method achieves SOTA performance, with +1.29 BLEU over five common languages and +2.93 BLEU on zero-shot languages, using only 10 hours of parallel speech per language. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_09232 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation Li, Xuanchen Cui, Chenrui Wang, Tianrui Ge, Meng Huang, Zikang Peng, Yizhou Li, Jin Lu, Yuheng Jiang, Yu Tashi, Nyima Wang, Longbiao Dang, Jianwu Computation and Language Sound Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation. However, existing approaches often overlook semantic commonalities across source languages, leading to biased translation performance. In this work, we propose POTSA (Parallel Optimal Transport for Speech Alignment), a new framework based on cross-lingual parallel speech pairs and Optimal Transport, designed to bridge high- and low-resource translation gaps. First, we introduce a Bias Compensation module to coarsely align initial speech representations. Second, we impose token-level OT constraints on a Q-Former using parallel pairs to establish fine-grained representation consistency. Then, we apply a layer scheduling strategy to focus OT constraints on semantically beneficial layers. Experiments on FLEURS show our method achieves SOTA performance, with +1.29 BLEU over five common languages and +2.93 BLEU on zero-shot languages, using only 10 hours of parallel speech per language. |
| title | POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation |
| topic | Computation and Language Sound |
| url | https://arxiv.org/abs/2511.09232 |