SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Haotian, Ouchi, Hiroki, Sakti, Sakriani
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914294229106688
author Tan, Haotian
Ouchi, Hiroki
Sakti, Sakriani
author_facet Tan, Haotian
Ouchi, Hiroki
Sakti, Sakriani
contents How to make human-interpreter-like read/write decisions for simultaneous speech translation (SimulST) systems? Current state-of-the-art systems formulate SimulST as a multi-turn dialogue task, requiring specialized interleaved training data and relying on computationally expensive large language model (LLM) inference for decision-making. In this paper, we propose SimulSense, a novel framework for SimulST that mimics human interpreters by continuously reading input speech and triggering write decisions to produce translation when a new sense unit is perceived. Experiments against two state-of-the-art baseline systems demonstrate that our proposed method achieves a superior quality-latency tradeoff and substantially improved real-time efficiency, where its decision-making is up to 9.6x faster than the baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21932
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation
Tan, Haotian
Ouchi, Hiroki
Sakti, Sakriani
Computation and Language
How to make human-interpreter-like read/write decisions for simultaneous speech translation (SimulST) systems? Current state-of-the-art systems formulate SimulST as a multi-turn dialogue task, requiring specialized interleaved training data and relying on computationally expensive large language model (LLM) inference for decision-making. In this paper, we propose SimulSense, a novel framework for SimulST that mimics human interpreters by continuously reading input speech and triggering write decisions to produce translation when a new sense unit is perceived. Experiments against two state-of-the-art baseline systems demonstrate that our proposed method achieves a superior quality-latency tradeoff and substantially improved real-time efficiency, where its decision-making is up to 9.6x faster than the baselines.
title SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation
topic Computation and Language
url https://arxiv.org/abs/2509.21932