Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Haoran, Sharaf, Amr, Chen, Yunmo, Tan, Weiting, Shen, Lingfeng, Van Durme, Benjamin, Murray, Kenton, Kim, Young Jin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917682526289920
author Xu, Haoran
Sharaf, Amr
Chen, Yunmo
Tan, Weiting
Shen, Lingfeng
Van Durme, Benjamin
Murray, Kenton
Kim, Young Jin
author_facet Xu, Haoran
Sharaf, Amr
Chen, Yunmo
Tan, Weiting
Shen, Lingfeng
Van Durme, Benjamin
Murray, Kenton
Kim, Young Jin
contents Moderate-sized large language models (LLMs) -- those with 7B or 13B parameters -- exhibit promising machine translation (MT) performance. However, even the top-performing 13B LLM-based translation models, like ALMA, does not match the performance of state-of-the-art conventional encoder-decoder translation models or larger-scale LLMs such as GPT-4. In this study, we bridge this performance gap. We first assess the shortcomings of supervised fine-tuning for LLMs in the MT task, emphasizing the quality issues present in the reference data, despite being human-generated. Then, in contrast to SFT which mimics reference translations, we introduce Contrastive Preference Optimization (CPO), a novel approach that trains models to avoid generating adequate but not perfect translations. Applying CPO to ALMA models with only 22K parallel sentences and 12M parameters yields significant improvements. The resulting model, called ALMA-R, can match or exceed the performance of the WMT competition winners and GPT-4 on WMT'21, WMT'22 and WMT'23 test datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08417
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Xu, Haoran
Sharaf, Amr
Chen, Yunmo
Tan, Weiting
Shen, Lingfeng
Van Durme, Benjamin
Murray, Kenton
Kim, Young Jin
Computation and Language
Moderate-sized large language models (LLMs) -- those with 7B or 13B parameters -- exhibit promising machine translation (MT) performance. However, even the top-performing 13B LLM-based translation models, like ALMA, does not match the performance of state-of-the-art conventional encoder-decoder translation models or larger-scale LLMs such as GPT-4. In this study, we bridge this performance gap. We first assess the shortcomings of supervised fine-tuning for LLMs in the MT task, emphasizing the quality issues present in the reference data, despite being human-generated. Then, in contrast to SFT which mimics reference translations, we introduce Contrastive Preference Optimization (CPO), a novel approach that trains models to avoid generating adequate but not perfect translations. Applying CPO to ALMA models with only 22K parallel sentences and 12M parameters yields significant improvements. The resulting model, called ALMA-R, can match or exceed the performance of the WMT competition winners and GPT-4 on WMT'21, WMT'22 and WMT'23 test datasets.
title Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2401.08417