Anticipating Future with Large Language Model for Simultaneous Machine Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ouyang, Siqi, Hrinchuk, Oleksii, Chen, Zhehuai, Lavrukhin, Vitaly, Balam, Jagadeesh, Li, Lei, Ginsburg, Boris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916769923334144
author Ouyang, Siqi
Hrinchuk, Oleksii
Chen, Zhehuai
Lavrukhin, Vitaly
Balam, Jagadeesh
Li, Lei
Ginsburg, Boris
author_facet Ouyang, Siqi
Hrinchuk, Oleksii
Chen, Zhehuai
Lavrukhin, Vitaly
Balam, Jagadeesh
Li, Lei
Ginsburg, Boris
contents Simultaneous machine translation (SMT) takes streaming input utterances and incrementally produces target text. Existing SMT methods mainly use the partial utterance that has already arrived at the input and the generated hypothesis. Motivated by human interpreters' technique to forecast future words before hearing them, we propose $\textbf{T}$ranslation by $\textbf{A}$nticipating $\textbf{F}$uture (TAF), a method to improve translation quality while retraining low latency. Its core idea is to use a large language model (LLM) to predict future source words and opportunistically translate without introducing too much risk. We evaluate our TAF and multiple baselines of SMT on four language directions. Experiments show that TAF achieves the best translation quality-latency trade-off and outperforms the baselines by up to 5 BLEU points at the same latency (three words). Code is released at https://github.com/owaski/TAF
format Preprint
id arxiv_https___arxiv_org_abs_2410_22499
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Anticipating Future with Large Language Model for Simultaneous Machine Translation
Ouyang, Siqi
Hrinchuk, Oleksii
Chen, Zhehuai
Lavrukhin, Vitaly
Balam, Jagadeesh
Li, Lei
Ginsburg, Boris
Computation and Language
Simultaneous machine translation (SMT) takes streaming input utterances and incrementally produces target text. Existing SMT methods mainly use the partial utterance that has already arrived at the input and the generated hypothesis. Motivated by human interpreters' technique to forecast future words before hearing them, we propose $\textbf{T}$ranslation by $\textbf{A}$nticipating $\textbf{F}$uture (TAF), a method to improve translation quality while retraining low latency. Its core idea is to use a large language model (LLM) to predict future source words and opportunistically translate without introducing too much risk. We evaluate our TAF and multiple baselines of SMT on four language directions. Experiments show that TAF achieves the best translation quality-latency trade-off and outperforms the baselines by up to 5 BLEU points at the same latency (three words). Code is released at https://github.com/owaski/TAF
title Anticipating Future with Large Language Model for Simultaneous Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2410.22499