CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909464649531392 |
|---|---|
| author | Cui, Guofeng Wang, Pichao Liu, Yang Ke, Zemian Liu, Zhu Bhat, Vimal |
| author_facet | Cui, Guofeng Wang, Pichao Liu, Yang Ke, Zemian Liu, Zhu Bhat, Vimal |
| contents | Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Optimization (DPO) has emerged as a simpler and more efficient alternative, but its performance depends heavily on the quality of preference data. To address this, we propose Confidence-Reward driven Preference Optimization (CRPO), a novel method that combines reward scores with model confidence to improve data selection for fine-tuning. CRPO selects challenging sentence pairs where the model is uncertain or underperforms, leading to more effective learning. While primarily designed for LLMs, CRPO also generalizes to encoder-decoder models like NLLB, demonstrating its versatility. Empirical results show that CRPO outperforms existing methods such as RS-DPO, RSO and MBR score in both translation accuracy and data efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_13927 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation Cui, Guofeng Wang, Pichao Liu, Yang Ke, Zemian Liu, Zhu Bhat, Vimal Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Optimization (DPO) has emerged as a simpler and more efficient alternative, but its performance depends heavily on the quality of preference data. To address this, we propose Confidence-Reward driven Preference Optimization (CRPO), a novel method that combines reward scores with model confidence to improve data selection for fine-tuning. CRPO selects challenging sentence pairs where the model is uncertain or underperforms, leading to more effective learning. While primarily designed for LLMs, CRPO also generalizes to encoder-decoder models like NLLB, demonstrating its versatility. Empirical results show that CRPO outperforms existing methods such as RS-DPO, RSO and MBR score in both translation accuracy and data efficiency. |
| title | CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2501.13927 |