Beyond Single-Reward: Multi-Pair, Multi-Perspective Preference Optimization for Machine Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hao, Xu, Linlong, Liu, Heng, Liu, Yangyang, Zhao, Xiaohu, Zeng, Bo, Shao, Liangying, Wang, Longyue, Luo, Weihua, Zhang, Kaifu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915555800252416
author Wang, Hao
Xu, Linlong
Liu, Heng
Liu, Yangyang
Zhao, Xiaohu
Zeng, Bo
Shao, Liangying
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
author_facet Wang, Hao
Xu, Linlong
Liu, Heng
Liu, Yangyang
Zhao, Xiaohu
Zeng, Bo
Shao, Liangying
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
contents Direct Preference Optimization (DPO) is a powerful paradigm for aligning Large Language Models (LLMs) to human preferences in Machine Translation (MT), but current methods are hindered by two fundamental challenges: (1) flawed reward signals from Quality Estimation (QE) models that overlook critical errors like translation hallucination, and (2) inefficient data utilization that discards valuable learning signals by selecting only a single win-loss pair. To address these limitations, we introduce M^2PO: Multi-Pair, Multi-Perspective Preference Optimization. Our framework integrates a multi-perspective reward engine that creates a more robust signal by combining two key viewpoints: a new hallucination penalty for factuality, and an innovative dynamic quality score that adaptively fuses external evaluations with the model's own evolving judgment. This is synergistically paired with a multi-pair construction strategy that systematically creates a comprehensive set of preference pairs from the entire pool of translation candidates. This synergistic approach ensures the model learns from a richer spectrum of quality trade-offs, leading to more robust and faithful translations. On challenging WMT21-22 benchmarks, M^2PO substantially outperforms existing preference optimization methods and demonstrates highly competitive performance against leading proprietary LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13434
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Single-Reward: Multi-Pair, Multi-Perspective Preference Optimization for Machine Translation
Wang, Hao
Xu, Linlong
Liu, Heng
Liu, Yangyang
Zhao, Xiaohu
Zeng, Bo
Shao, Liangying
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
Computation and Language
Direct Preference Optimization (DPO) is a powerful paradigm for aligning Large Language Models (LLMs) to human preferences in Machine Translation (MT), but current methods are hindered by two fundamental challenges: (1) flawed reward signals from Quality Estimation (QE) models that overlook critical errors like translation hallucination, and (2) inefficient data utilization that discards valuable learning signals by selecting only a single win-loss pair. To address these limitations, we introduce M^2PO: Multi-Pair, Multi-Perspective Preference Optimization. Our framework integrates a multi-perspective reward engine that creates a more robust signal by combining two key viewpoints: a new hallucination penalty for factuality, and an innovative dynamic quality score that adaptively fuses external evaluations with the model's own evolving judgment. This is synergistically paired with a multi-pair construction strategy that systematically creates a comprehensive set of preference pairs from the entire pool of translation candidates. This synergistic approach ensures the model learns from a richer spectrum of quality trade-offs, leading to more robust and faithful translations. On challenging WMT21-22 benchmarks, M^2PO substantially outperforms existing preference optimization methods and demonstrates highly competitive performance against leading proprietary LLMs.
title Beyond Single-Reward: Multi-Pair, Multi-Perspective Preference Optimization for Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2510.13434