Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Jian, Yin, Shenglin, Zhang, Yujia, Zhao, Alan, Chen, Xi, Zhou, Xiaohui, Xu, Pengfei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914174031888384
author Li, Jian
Yin, Shenglin
Zhang, Yujia
Zhao, Alan
Chen, Xi
Zhou, Xiaohui
Xu, Pengfei
author_facet Li, Jian
Yin, Shenglin
Zhang, Yujia
Zhao, Alan
Chen, Xi
Zhou, Xiaohui
Xu, Pengfei
contents Direct Preference Optimization (DPO) is a widely used reinforcement learning from human feedback (RLHF) method across various domains. Recent research has increasingly focused on the role of token importance in improving DPO effectiveness. It is observed that identical or semantically similar content (defined as ambiguous content) frequently appears within the preference pairs. We hypothesize that the presence of ambiguous content during DPO training may introduce ambiguity, thereby limiting further improvements in alignment. Through mathematical analysis and proof-of-concept experiments, we reveal that ambiguous content may potentially introduce ambiguities, thereby degrading performance. To address this issue, we introduce Ambiguity Awareness Optimization (AAO), a simple yet effective approach that automatically re-weights ambiguous content to reduce ambiguities by calculating semantic similarity from preference pairs. Through extensive experiments, we demonstrate that AAO consistently and significantly surpasses state-of-the-art approaches in performance, without markedly increasing response length, across multiple model scales and widely adopted benchmark datasets, including AlpacaEval 2, MT-Bench, and Arena-Hard. Specifically, AAO outperforms DPO by up to 8.9 points on AlpacaEval 2 and achieves an improvement of by up to 15.0 points on Arena-Hard.
format Preprint
id arxiv_https___arxiv_org_abs_2511_23391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
Li, Jian
Yin, Shenglin
Zhang, Yujia
Zhao, Alan
Chen, Xi
Zhou, Xiaohui
Xu, Pengfei
Computation and Language
Direct Preference Optimization (DPO) is a widely used reinforcement learning from human feedback (RLHF) method across various domains. Recent research has increasingly focused on the role of token importance in improving DPO effectiveness. It is observed that identical or semantically similar content (defined as ambiguous content) frequently appears within the preference pairs. We hypothesize that the presence of ambiguous content during DPO training may introduce ambiguity, thereby limiting further improvements in alignment. Through mathematical analysis and proof-of-concept experiments, we reveal that ambiguous content may potentially introduce ambiguities, thereby degrading performance. To address this issue, we introduce Ambiguity Awareness Optimization (AAO), a simple yet effective approach that automatically re-weights ambiguous content to reduce ambiguities by calculating semantic similarity from preference pairs. Through extensive experiments, we demonstrate that AAO consistently and significantly surpasses state-of-the-art approaches in performance, without markedly increasing response length, across multiple model scales and widely adopted benchmark datasets, including AlpacaEval 2, MT-Bench, and Arena-Hard. Specifically, AAO outperforms DPO by up to 8.9 points on AlpacaEval 2 and achieves an improvement of by up to 15.0 points on Arena-Hard.
title Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
topic Computation and Language
url https://arxiv.org/abs/2511.23391