Mufu: Multilingual Fused Learning for Low-Resource Translation with LLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lim, Zheng Wei, Gupta, Nitish, Yu, Honglin, Cohn, Trevor
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912696635490304
author Lim, Zheng Wei
Gupta, Nitish
Yu, Honglin
Cohn, Trevor
author_facet Lim, Zheng Wei
Gupta, Nitish
Yu, Honglin
Cohn, Trevor
contents Multilingual large language models (LLMs) are great translators, but this is largely limited to high-resource languages. For many LLMs, translating in and out of low-resource languages remains a challenging task. To maximize data efficiency in this low-resource setting, we introduce Mufu, which includes a selection of automatically generated multilingual candidates and an instruction to correct inaccurate translations in the prompt. Mufu prompts turn a translation task into a postediting one, and seek to harness the LLM's reasoning capability with auxiliary translation candidates, from which the model is required to assess the input quality, align the semantics cross-lingually, copy from relevant inputs and override instances that are incorrect. Our experiments on En-XX translations over the Flores-200 dataset show LLMs finetuned against Mufu-style prompts are robust to poor quality auxiliary translation candidates, achieving performance superior to NLLB 1.3B distilled model in 64% of low- and very-low-resource language pairs. We then distill these models to reduce inference cost, while maintaining on average 3.1 chrF improvement over finetune-only baseline in low-resource translations.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13949
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mufu: Multilingual Fused Learning for Low-Resource Translation with LLM
Lim, Zheng Wei
Gupta, Nitish
Yu, Honglin
Cohn, Trevor
Computation and Language
Multilingual large language models (LLMs) are great translators, but this is largely limited to high-resource languages. For many LLMs, translating in and out of low-resource languages remains a challenging task. To maximize data efficiency in this low-resource setting, we introduce Mufu, which includes a selection of automatically generated multilingual candidates and an instruction to correct inaccurate translations in the prompt. Mufu prompts turn a translation task into a postediting one, and seek to harness the LLM's reasoning capability with auxiliary translation candidates, from which the model is required to assess the input quality, align the semantics cross-lingually, copy from relevant inputs and override instances that are incorrect. Our experiments on En-XX translations over the Flores-200 dataset show LLMs finetuned against Mufu-style prompts are robust to poor quality auxiliary translation candidates, achieving performance superior to NLLB 1.3B distilled model in 64% of low- and very-low-resource language pairs. We then distill these models to reduce inference cost, while maintaining on average 3.1 chrF improvement over finetune-only baseline in low-resource translations.
title Mufu: Multilingual Fused Learning for Low-Resource Translation with LLM
topic Computation and Language
url https://arxiv.org/abs/2409.13949