Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bouthors, Maxime, Crego, Josep, Yvon, François
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912619653234688
author Bouthors, Maxime
Crego, Josep
Yvon, François
author_facet Bouthors, Maxime
Crego, Josep
Yvon, François
contents Conventional retrieval-augmented neural machine translation (RANMT) systems leverage bilingual corpora, e.g., translation memories (TMs). Yet, in many settings, monolingual corpora in the target language are often available. This work explores ways to take advantage of such resources by directly retrieving relevant target language segments, based on a source-side query. For this, we design improved cross-lingual retrieval systems, trained with both sentence level and word-level matching objectives. In our experiments with three RANMT architectures, we assess such cross-lingual objectives in a controlled setting, reaching performances that match those of standard TM-based models. We also showcase our method on a real-world settings, using much larger monolingual and observe strong improvements over both the baseline setting and general-purpose cross-lingual retrievers.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21747
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
Bouthors, Maxime
Crego, Josep
Yvon, François
Computation and Language
I.2.7
Conventional retrieval-augmented neural machine translation (RANMT) systems leverage bilingual corpora, e.g., translation memories (TMs). Yet, in many settings, monolingual corpora in the target language are often available. This work explores ways to take advantage of such resources by directly retrieving relevant target language segments, based on a source-side query. For this, we design improved cross-lingual retrieval systems, trained with both sentence level and word-level matching objectives. In our experiments with three RANMT architectures, we assess such cross-lingual objectives in a controlled setting, reaching performances that match those of standard TM-based models. We also showcase our method on a real-world settings, using much larger monolingual and observe strong improvements over both the baseline setting and general-purpose cross-lingual retrievers.
title Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2504.21747