Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Lijun, Li, Lin, Qi, Yajie, Song, Huizhong, Yang, Yaodong, Wang, Jun, Wei, Wei
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2505.20359
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909626856898560
author Zhang, Lijun
Li, Lin
Qi, Yajie
Song, Huizhong
Yang, Yaodong
Wang, Jun
Wei, Wei
author_facet Zhang, Lijun
Li, Lin
Qi, Yajie
Song, Huizhong
Yang, Yaodong
Wang, Jun
Wei, Wei
contents When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also introduces potential risks due to deviations from the reference model's intended behavior. Most existing methods typically introduce KL divergence to constrain deviations between the trained model and the reference model; however, this may not be sufficient in certain applications that require tight risk control. In this paper, we introduce Risk-aware Direct Preference Optimization (Ra-DPO), a novel approach that incorporates risk-awareness by employing a class of nested risk measures. This approach formulates a constrained risk-aware advantage function maximization problem and then converts the Bradley-Terry model into a token-level representation. The objective function maximizes the likelihood of the policy while suppressing the deviation between a trained model and the reference model using a sequential risk ratio, thereby enhancing the model's risk-awareness. Experimental results across three open-source datasets: IMDb Dataset, Anthropic HH Dataset, and AlpacaEval, demonstrate the proposed method's superior performance in balancing alignment performance and model drift. Our code is opensourced at https://github.com/zlj123-max/Ra-DPO.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20359
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Risk-aware Direct Preference Optimization under Nested Risk Measure
Zhang, Lijun
Li, Lin
Qi, Yajie
Song, Huizhong
Yang, Yaodong
Wang, Jun
Wei, Wei
Machine Learning
Artificial Intelligence
When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also introduces potential risks due to deviations from the reference model's intended behavior. Most existing methods typically introduce KL divergence to constrain deviations between the trained model and the reference model; however, this may not be sufficient in certain applications that require tight risk control. In this paper, we introduce Risk-aware Direct Preference Optimization (Ra-DPO), a novel approach that incorporates risk-awareness by employing a class of nested risk measures. This approach formulates a constrained risk-aware advantage function maximization problem and then converts the Bradley-Terry model into a token-level representation. The objective function maximizes the likelihood of the policy while suppressing the deviation between a trained model and the reference model using a sequential risk ratio, thereby enhancing the model's risk-awareness. Experimental results across three open-source datasets: IMDb Dataset, Anthropic HH Dataset, and AlpacaEval, demonstrate the proposed method's superior performance in balancing alignment performance and model drift. Our code is opensourced at https://github.com/zlj123-max/Ra-DPO.
title Risk-aware Direct Preference Optimization under Nested Risk Measure
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.20359