InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ouyang, Siqi, Xu, Xi, Li, Lei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913894784565248
author Ouyang, Siqi
Xu, Xi
Li, Lei
author_facet Ouyang, Siqi
Xu, Xi
Li, Lei
contents Simultaneous translation of unbounded streaming speech remains a challenging problem due to the need for effectively processing the history speech context and past translations so that quality and latency, including computation overhead, can be balanced. Most prior works assume pre-segmented speech, limiting their real-world applicability. In this paper, we propose InfiniSST, a novel approach that formulates SST as a multi-turn dialogue task, enabling seamless translation of unbounded speech. We construct translation trajectories and robust segments from MuST-C with multi-latency augmentation during training and develop a key-value (KV) cache management strategy to facilitate efficient inference. Experiments on MuST-C En-Es, En-De, and En-Zh demonstrate that InfiniSST reduces computation-aware latency by 0.5 to 1 second while maintaining the same translation quality compared to baselines. Ablation studies further validate the contributions of our data construction and cache management strategy. We release the code and demo at https://github.com/LeiLiLab/InfiniSST
format Preprint
id arxiv_https___arxiv_org_abs_2503_02969
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model
Ouyang, Siqi
Xu, Xi
Li, Lei
Computation and Language
Artificial Intelligence
Simultaneous translation of unbounded streaming speech remains a challenging problem due to the need for effectively processing the history speech context and past translations so that quality and latency, including computation overhead, can be balanced. Most prior works assume pre-segmented speech, limiting their real-world applicability. In this paper, we propose InfiniSST, a novel approach that formulates SST as a multi-turn dialogue task, enabling seamless translation of unbounded speech. We construct translation trajectories and robust segments from MuST-C with multi-latency augmentation during training and develop a key-value (KV) cache management strategy to facilitate efficient inference. Experiments on MuST-C En-Es, En-De, and En-Zh demonstrate that InfiniSST reduces computation-aware latency by 0.5 to 1 second while maintaining the same translation quality compared to baselines. Ablation studies further validate the contributions of our data construction and cache management strategy. We release the code and demo at https://github.com/LeiLiLab/InfiniSST
title InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.02969