SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910057231286272 |
|---|---|
| author | Djanibekov, Amirbek Bentivogli, Luisa Negri, Matteo Papi, Sara |
| author_facet | Djanibekov, Amirbek Bentivogli, Luisa Negri, Matteo Papi, Sara |
| contents | Simultaneous speech-to-speech translation (SimulS2S) is essential for real-time multilingual communication, with increasing integration into meeting and streaming platforms. Despite this, SimulS2S remains underexplored in research, where current solutions often rely on resource-intensive training procedures and operate on short-form, pre-segmented utterances, failing to generalize to continuous speech. To bridge this gap, we propose SimulU, the first training-free policy for long-form SimulS2S. SimulU adopts history management and speech output selection strategies that exploit cross-attention in pre-trained end-to-end models to regulate both input history and output generation. Evaluations on MuST-C across 8 languages show that SimulU achieves a better or comparable quality-latency trade-off against strong cascaded models. By eliminating the need for ad-hoc training, SimulU offers a promising path to end-to-end SimulS2S in realistic, long-form scenarios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_16924 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation Djanibekov, Amirbek Bentivogli, Luisa Negri, Matteo Papi, Sara Audio and Speech Processing Artificial Intelligence Computation and Language Simultaneous speech-to-speech translation (SimulS2S) is essential for real-time multilingual communication, with increasing integration into meeting and streaming platforms. Despite this, SimulS2S remains underexplored in research, where current solutions often rely on resource-intensive training procedures and operate on short-form, pre-segmented utterances, failing to generalize to continuous speech. To bridge this gap, we propose SimulU, the first training-free policy for long-form SimulS2S. SimulU adopts history management and speech output selection strategies that exploit cross-attention in pre-trained end-to-end models to regulate both input history and output generation. Evaluations on MuST-C across 8 languages show that SimulU achieves a better or comparable quality-latency trade-off against strong cascaded models. By eliminating the need for ad-hoc training, SimulU offers a promising path to end-to-end SimulS2S in realistic, long-form scenarios. |
| title | SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2603.16924 |