Weights-Rotated Preference Optimization for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chenxu, Jia, Ruipeng, Zheng, Mingyu, Gu, Naibin, Lin, Zheng, Chen, Siyuan, Yin, Weichong, Wu, Hua, Wang, Weiping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909751452893184
author Yang, Chenxu
Jia, Ruipeng
Zheng, Mingyu
Gu, Naibin
Lin, Zheng
Chen, Siyuan
Yin, Weichong
Wu, Hua
Wang, Weiping
author_facet Yang, Chenxu
Jia, Ruipeng
Zheng, Mingyu
Gu, Naibin
Lin, Zheng
Chen, Siyuan
Yin, Weichong
Wu, Hua
Wang, Weiping
contents Despite the efficacy of Direct Preference Optimization (DPO) in aligning Large Language Models (LLMs), reward hacking remains a pivotal challenge. This issue emerges when LLMs excessively reduce the probability of rejected completions to achieve high rewards, without genuinely meeting their intended goals. As a result, this leads to overly lengthy generation lacking diversity, as well as catastrophic forgetting of knowledge. We investigate the underlying reason behind this issue, which is representation redundancy caused by neuron collapse in the parameter space. Hence, we propose a novel Weights-Rotated Preference Optimization (RoPO) algorithm, which implicitly constrains the output layer logits with the KL divergence inherited from DPO and explicitly constrains the intermediate hidden states by fine-tuning on a multi-granularity orthogonal matrix. This design prevents the policy model from deviating too far from the reference model, thereby retaining the knowledge and expressive capabilities acquired during pre-training and SFT stages. Our RoPO achieves up to a 3.27-point improvement on AlpacaEval 2, and surpasses the best baseline by 6.2 to 7.5 points on MT-Bench with merely 0.015% of the trainable parameters, demonstrating its effectiveness in alleviating the reward hacking problem of DPO.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17637
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Weights-Rotated Preference Optimization for Large Language Models
Yang, Chenxu
Jia, Ruipeng
Zheng, Mingyu
Gu, Naibin
Lin, Zheng
Chen, Siyuan
Yin, Weichong
Wu, Hua
Wang, Weiping
Computation and Language
Artificial Intelligence
Despite the efficacy of Direct Preference Optimization (DPO) in aligning Large Language Models (LLMs), reward hacking remains a pivotal challenge. This issue emerges when LLMs excessively reduce the probability of rejected completions to achieve high rewards, without genuinely meeting their intended goals. As a result, this leads to overly lengthy generation lacking diversity, as well as catastrophic forgetting of knowledge. We investigate the underlying reason behind this issue, which is representation redundancy caused by neuron collapse in the parameter space. Hence, we propose a novel Weights-Rotated Preference Optimization (RoPO) algorithm, which implicitly constrains the output layer logits with the KL divergence inherited from DPO and explicitly constrains the intermediate hidden states by fine-tuning on a multi-granularity orthogonal matrix. This design prevents the policy model from deviating too far from the reference model, thereby retaining the knowledge and expressive capabilities acquired during pre-training and SFT stages. Our RoPO achieves up to a 3.27-point improvement on AlpacaEval 2, and surpasses the best baseline by 6.2 to 7.5 points on MT-Bench with merely 0.015% of the trainable parameters, demonstrating its effectiveness in alleviating the reward hacking problem of DPO.
title Weights-Rotated Preference Optimization for Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.17637