KL Penalty Control via Perturbation for Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Sangkyu, Han, Janghoon, Song, Hosung, Choi, Stanley Jungkyu, Lee, Honglak, Yu, Youngjae
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909868718292992
author Lee, Sangkyu
Han, Janghoon
Song, Hosung
Choi, Stanley Jungkyu
Lee, Honglak
Yu, Youngjae
author_facet Lee, Sangkyu
Han, Janghoon
Song, Hosung
Choi, Stanley Jungkyu
Lee, Honglak
Yu, Youngjae
contents Direct Preference Optimization (DPO) demonstrates the advantage of aligning a large language model with human preference using only an offline dataset. However, DPO has the limitation that the KL penalty, which prevents excessive deviation from the reference model, is static throughout the training process. Several methods claim to change this static KL penalty of DPO into a dynamic one, but no approach can adaptively assign different KL penalties for each preference pair. In this paper, we propose $\varepsilon$-Direct Preference Optimization ($\varepsilon$-DPO), which allows adaptive control of the KL penalty strength $β$ for each preference pair. Specifically, $\varepsilon$-DPO adaptively controls $β$ for each preference pair based on the monotonicity of logits as a preference model under the perturbation of $β$ during training. This is equivalent to adjusting the KL penalty by checking whether the change in training-time temperature can lead to better preference confidence as preference models by simply reusing the logit of the current policy and the reference policy. Experimental results show that the simple criterion of $\varepsilon$-DPO for KL penalty relaxation significantly improves DPO compared to most existing direct alignment algorithms on general chatbot benchmarks and reveal that this KL penalty control criterion can reflect confusion as a preference model and provide an efficient KL trade-off, highlighting the significance of instance-level adaptive KL penalty control in DPO.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13177
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KL Penalty Control via Perturbation for Direct Preference Optimization
Lee, Sangkyu
Han, Janghoon
Song, Hosung
Choi, Stanley Jungkyu
Lee, Honglak
Yu, Youngjae
Machine Learning
Artificial Intelligence
Direct Preference Optimization (DPO) demonstrates the advantage of aligning a large language model with human preference using only an offline dataset. However, DPO has the limitation that the KL penalty, which prevents excessive deviation from the reference model, is static throughout the training process. Several methods claim to change this static KL penalty of DPO into a dynamic one, but no approach can adaptively assign different KL penalties for each preference pair. In this paper, we propose $\varepsilon$-Direct Preference Optimization ($\varepsilon$-DPO), which allows adaptive control of the KL penalty strength $β$ for each preference pair. Specifically, $\varepsilon$-DPO adaptively controls $β$ for each preference pair based on the monotonicity of logits as a preference model under the perturbation of $β$ during training. This is equivalent to adjusting the KL penalty by checking whether the change in training-time temperature can lead to better preference confidence as preference models by simply reusing the logit of the current policy and the reference policy. Experimental results show that the simple criterion of $\varepsilon$-DPO for KL penalty relaxation significantly improves DPO compared to most existing direct alignment algorithms on general chatbot benchmarks and reveal that this KL penalty control criterion can reflect confusion as a preference model and provide an efficient KL trade-off, highlighting the significance of instance-level adaptive KL penalty control in DPO.
title KL Penalty Control via Perturbation for Direct Preference Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.13177