SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914294229106688 |
|---|---|
| author | Tan, Haotian Ouchi, Hiroki Sakti, Sakriani |
| author_facet | Tan, Haotian Ouchi, Hiroki Sakti, Sakriani |
| contents | How to make human-interpreter-like read/write decisions for simultaneous speech translation (SimulST) systems? Current state-of-the-art systems formulate SimulST as a multi-turn dialogue task, requiring specialized interleaved training data and relying on computationally expensive large language model (LLM) inference for decision-making. In this paper, we propose SimulSense, a novel framework for SimulST that mimics human interpreters by continuously reading input speech and triggering write decisions to produce translation when a new sense unit is perceived. Experiments against two state-of-the-art baseline systems demonstrate that our proposed method achieves a superior quality-latency tradeoff and substantially improved real-time efficiency, where its decision-making is up to 9.6x faster than the baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_21932 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation Tan, Haotian Ouchi, Hiroki Sakti, Sakriani Computation and Language How to make human-interpreter-like read/write decisions for simultaneous speech translation (SimulST) systems? Current state-of-the-art systems formulate SimulST as a multi-turn dialogue task, requiring specialized interleaved training data and relying on computationally expensive large language model (LLM) inference for decision-making. In this paper, we propose SimulSense, a novel framework for SimulST that mimics human interpreters by continuously reading input speech and triggering write decisions to produce translation when a new sense unit is perceived. Experiments against two state-of-the-art baseline systems demonstrate that our proposed method achieves a superior quality-latency tradeoff and substantially improved real-time efficiency, where its decision-making is up to 9.6x faster than the baselines. |
| title | SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.21932 |