CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cui, Guofeng, Wang, Pichao, Liu, Yang, Ke, Zemian, Liu, Zhu, Bhat, Vimal
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909464649531392
author Cui, Guofeng
Wang, Pichao
Liu, Yang
Ke, Zemian
Liu, Zhu
Bhat, Vimal
author_facet Cui, Guofeng
Wang, Pichao
Liu, Yang
Ke, Zemian
Liu, Zhu
Bhat, Vimal
contents Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Optimization (DPO) has emerged as a simpler and more efficient alternative, but its performance depends heavily on the quality of preference data. To address this, we propose Confidence-Reward driven Preference Optimization (CRPO), a novel method that combines reward scores with model confidence to improve data selection for fine-tuning. CRPO selects challenging sentence pairs where the model is uncertain or underperforms, leading to more effective learning. While primarily designed for LLMs, CRPO also generalizes to encoder-decoder models like NLLB, demonstrating its versatility. Empirical results show that CRPO outperforms existing methods such as RS-DPO, RSO and MBR score in both translation accuracy and data efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13927
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation
Cui, Guofeng
Wang, Pichao
Liu, Yang
Ke, Zemian
Liu, Zhu
Bhat, Vimal
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Optimization (DPO) has emerged as a simpler and more efficient alternative, but its performance depends heavily on the quality of preference data. To address this, we propose Confidence-Reward driven Preference Optimization (CRPO), a novel method that combines reward scores with model confidence to improve data selection for fine-tuning. CRPO selects challenging sentence pairs where the model is uncertain or underperforms, leading to more effective learning. While primarily designed for LLMs, CRPO also generalizes to encoder-decoder models like NLLB, demonstrating its versatility. Empirical results show that CRPO outperforms existing methods such as RS-DPO, RSO and MBR score in both translation accuracy and data efficiency.
title CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.13927