Saved in:
Bibliographic Details
Main Authors: Ahmed, Zeeshan, Seide, Frank, Moritz, Niko, Lin, Ju, Xie, Ruiming, Merello, Simone, Liu, Zhe, Fuegen, Christian
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.13358
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913996566691840
author Ahmed, Zeeshan
Seide, Frank
Moritz, Niko
Lin, Ju
Xie, Ruiming
Merello, Simone
Liu, Zhe
Fuegen, Christian
author_facet Ahmed, Zeeshan
Seide, Frank
Moritz, Niko
Lin, Ju
Xie, Ruiming
Merello, Simone
Liu, Zhe
Fuegen, Christian
contents This paper tackles several challenges that arise when integrating Automatic Speech Recognition (ASR) and Machine Translation (MT) for real-time, on-device streaming speech translation. Although state-of-the-art ASR systems based on Recurrent Neural Network Transducers (RNN-T) can perform real-time transcription, achieving streaming translation in real-time remains a significant challenge. To address this issue, we propose a simultaneous translation approach that effectively balances translation quality and latency. We also investigate efficient integration of ASR and MT, leveraging linguistic cues generated by the ASR system to manage context and utilizing efficient beam-search pruning techniques such as time-out and forced finalization to maintain system's real-time factor. We apply our approach to an on-device bilingual conversational speech translation and demonstrate that our techniques outperform baselines in terms of latency and quality. Notably, our technique narrows the quality gap with non-streaming translation systems, paving the way for more accurate and efficient real-time speech translation.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Overcoming Latency Bottlenecks in On-Device Speech Translation: A Cascaded Approach with Alignment-Based Streaming MT
Ahmed, Zeeshan
Seide, Frank
Moritz, Niko
Lin, Ju
Xie, Ruiming
Merello, Simone
Liu, Zhe
Fuegen, Christian
Computation and Language
Artificial Intelligence
This paper tackles several challenges that arise when integrating Automatic Speech Recognition (ASR) and Machine Translation (MT) for real-time, on-device streaming speech translation. Although state-of-the-art ASR systems based on Recurrent Neural Network Transducers (RNN-T) can perform real-time transcription, achieving streaming translation in real-time remains a significant challenge. To address this issue, we propose a simultaneous translation approach that effectively balances translation quality and latency. We also investigate efficient integration of ASR and MT, leveraging linguistic cues generated by the ASR system to manage context and utilizing efficient beam-search pruning techniques such as time-out and forced finalization to maintain system's real-time factor. We apply our approach to an on-device bilingual conversational speech translation and demonstrate that our techniques outperform baselines in terms of latency and quality. Notably, our technique narrows the quality gap with non-streaming translation systems, paving the way for more accurate and efficient real-time speech translation.
title Overcoming Latency Bottlenecks in On-Device Speech Translation: A Cascaded Approach with Alignment-Based Streaming MT
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.13358