Extending Multilingual Machine Translation through Imitation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lai, Wen, Hangya, Viktor, Shen, Yingli, Fraser, Alexander
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914174141988864
author Lai, Wen
Hangya, Viktor
Shen, Yingli
Fraser, Alexander
author_facet Lai, Wen
Hangya, Viktor
Shen, Yingli
Fraser, Alexander
contents Despite the growing variety of languages supported by existing multilingual neural machine translation (MNMT) models, most of the world's languages are still being left behind. We aim to extend large-scale MNMT models to incorporate a new language, enabling translations between this new language and all previously supported languages, even in the challenging scenario where only a parallel corpus between the new language and English is available. Previous methods, such as continued training on parallel data including the new language, often suffer from catastrophic forgetting, which degrades performance on other languages. We propose a novel approach Imit-MNMT which treats this task as an imitation learning problem, a technique widely used in computer vision but less explored in natural language processing. Specifically, we leverage an expert model to generate pseudo-parallel corpora between the new language and the existing languages. We then introduce a data distribution imitation strategy using language-specific weighting, alongside a translation behavior imitation mechanism. Extensive experiments show that our approach significantly improves translation performance between the new and existing languages while mitigating catastrophic forgetting.
format Preprint
id arxiv_https___arxiv_org_abs_2311_08538
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Extending Multilingual Machine Translation through Imitation Learning
Lai, Wen
Hangya, Viktor
Shen, Yingli
Fraser, Alexander
Computation and Language
Despite the growing variety of languages supported by existing multilingual neural machine translation (MNMT) models, most of the world's languages are still being left behind. We aim to extend large-scale MNMT models to incorporate a new language, enabling translations between this new language and all previously supported languages, even in the challenging scenario where only a parallel corpus between the new language and English is available. Previous methods, such as continued training on parallel data including the new language, often suffer from catastrophic forgetting, which degrades performance on other languages. We propose a novel approach Imit-MNMT which treats this task as an imitation learning problem, a technique widely used in computer vision but less explored in natural language processing. Specifically, we leverage an expert model to generate pseudo-parallel corpora between the new language and the existing languages. We then introduce a data distribution imitation strategy using language-specific weighting, alongside a translation behavior imitation mechanism. Extensive experiments show that our approach significantly improves translation performance between the new and existing languages while mitigating catastrophic forgetting.
title Extending Multilingual Machine Translation through Imitation Learning
topic Computation and Language
url https://arxiv.org/abs/2311.08538