Towards Boosting Many-to-Many Multilingual Machine Translation with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Pengzhi, He, Zhongjun, Wu, Hua, Wang, Haifeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910320578002944
author Gao, Pengzhi
He, Zhongjun
Wu, Hua
Wang, Haifeng
author_facet Gao, Pengzhi
He, Zhongjun
Wu, Hua
Wang, Haifeng
contents The training paradigm for machine translation has gradually shifted, from learning neural machine translation (NMT) models with extensive parallel corpora to instruction finetuning on multilingual large language models (LLMs) with high-quality translation pairs. In this paper, we focus on boosting many-to-many multilingual translation of LLMs with an emphasis on zero-shot translation directions. We demonstrate that prompt strategies adopted during finetuning are crucial to zero-shot translation and introduce a cross-lingual consistency regularization, XConST, to bridge the representation gap among different languages and improve zero-shot translation performance. XConST is not a new method, but a version of CrossConST (Gao et al., 2023a) adapted for translation instruction finetuning with LLMs. Experimental results on ALMA (Xu et al., 2023), Tower (Team, 2024), and LLaMA-2 (Touvron et al., 2023) show that our approach consistently improves translation performance. Our implementations are available at https://github.com/gpengzhi/CrossConST-LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2401_05861
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Boosting Many-to-Many Multilingual Machine Translation with Large Language Models
Gao, Pengzhi
He, Zhongjun
Wu, Hua
Wang, Haifeng
Computation and Language
The training paradigm for machine translation has gradually shifted, from learning neural machine translation (NMT) models with extensive parallel corpora to instruction finetuning on multilingual large language models (LLMs) with high-quality translation pairs. In this paper, we focus on boosting many-to-many multilingual translation of LLMs with an emphasis on zero-shot translation directions. We demonstrate that prompt strategies adopted during finetuning are crucial to zero-shot translation and introduce a cross-lingual consistency regularization, XConST, to bridge the representation gap among different languages and improve zero-shot translation performance. XConST is not a new method, but a version of CrossConST (Gao et al., 2023a) adapted for translation instruction finetuning with LLMs. Experimental results on ALMA (Xu et al., 2023), Tower (Team, 2024), and LLaMA-2 (Touvron et al., 2023) show that our approach consistently improves translation performance. Our implementations are available at https://github.com/gpengzhi/CrossConST-LLM.
title Towards Boosting Many-to-Many Multilingual Machine Translation with Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2401.05861