Adaptive Preference Optimization with Uncertainty-aware Utility Anchor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiaobo, Jia, Zixia, Li, Jiaqi, Liu, Qi, Zheng, Zilong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917527389470720
author Wang, Xiaobo
Jia, Zixia
Li, Jiaqi
Liu, Qi
Zheng, Zilong
author_facet Wang, Xiaobo
Jia, Zixia
Li, Jiaqi
Liu, Qi
Zheng, Zilong
contents Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning, one of the most popular approaches, stands out for its efficiency in reward modeling. However, these methods typically follow the convention to use Bradley-Terry (BT) reward modeling that faces several critical assumptions, including the requirement for pairwise training data, model distribution shifting, human rationality assumption, etc. To address these limitations, we propose a general framework for offline preference optimization methods, Adaptive Preference Optimization with Utility Anchor (UAPO), which introduces an anchoring function to estimate the uncertainties brought from preference data annotation. Our method enables training even in scenarios where the data is unpaired, significantly enhancing data utilization efficiency. Moreover, the anchor design makes UAPO more robust in the training process. Experimental results demonstrate that UAPO achieves competitive outcomes without the strict dependency on data pairing, paving the way for more flexible and effective preference optimization methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10515
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
Wang, Xiaobo
Jia, Zixia
Li, Jiaqi
Liu, Qi
Zheng, Zilong
Machine Learning
Computation and Language
Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning, one of the most popular approaches, stands out for its efficiency in reward modeling. However, these methods typically follow the convention to use Bradley-Terry (BT) reward modeling that faces several critical assumptions, including the requirement for pairwise training data, model distribution shifting, human rationality assumption, etc. To address these limitations, we propose a general framework for offline preference optimization methods, Adaptive Preference Optimization with Utility Anchor (UAPO), which introduces an anchoring function to estimate the uncertainties brought from preference data annotation. Our method enables training even in scenarios where the data is unpaired, significantly enhancing data utilization efficiency. Moreover, the anchor design makes UAPO more robust in the training process. Experimental results demonstrate that UAPO achieves competitive outcomes without the strict dependency on data pairing, paving the way for more flexible and effective preference optimization methods.
title Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.10515