Reverse Preference Optimization for Complex Instruction Following

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Xiang, Lin, Ting-En, Fang, Feiteng, Wu, Yuchuan, Li, Hangyu, Qu, Yuzhong, Huang, Fei, Li, Yongbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915309856751616
author Huang, Xiang
Lin, Ting-En
Fang, Feiteng
Wu, Yuchuan
Li, Hangyu
Qu, Yuzhong
Huang, Fei
Li, Yongbin
author_facet Huang, Xiang
Lin, Ting-En
Fang, Feiteng
Wu, Yuchuan
Li, Hangyu
Qu, Yuzhong
Huang, Fei
Li, Yongbin
contents Instruction following (IF) is a critical capability for large language models (LLMs). However, handling complex instructions with multiple constraints remains challenging. Previous methods typically select preference pairs based on the number of constraints they satisfy, introducing noise where chosen examples may fail to follow some constraints and rejected examples may excel in certain respects over the chosen ones. To address the challenge of aligning with multiple preferences, we propose a simple yet effective method called Reverse Preference Optimization (RPO). It mitigates noise in preference pairs by dynamically reversing the constraints within the instruction to ensure the chosen response is perfect, alleviating the burden of extensive sampling and filtering to collect perfect responses. Besides, reversal also enlarges the gap between chosen and rejected responses, thereby clarifying the optimization direction and making it more robust to noise. We evaluate RPO on two multi-turn IF benchmarks, Sysbench and Multi-IF, demonstrating average improvements over the DPO baseline of 4.6 and 2.5 points (on Llama-3.1 8B), respectively. Moreover, RPO scales effectively across model sizes (8B to 70B parameters), with the 70B RPO model surpassing GPT-4o.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22172
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reverse Preference Optimization for Complex Instruction Following
Huang, Xiang
Lin, Ting-En
Fang, Feiteng
Wu, Yuchuan
Li, Hangyu
Qu, Yuzhong
Huang, Fei
Li, Yongbin
Computation and Language
Instruction following (IF) is a critical capability for large language models (LLMs). However, handling complex instructions with multiple constraints remains challenging. Previous methods typically select preference pairs based on the number of constraints they satisfy, introducing noise where chosen examples may fail to follow some constraints and rejected examples may excel in certain respects over the chosen ones. To address the challenge of aligning with multiple preferences, we propose a simple yet effective method called Reverse Preference Optimization (RPO). It mitigates noise in preference pairs by dynamically reversing the constraints within the instruction to ensure the chosen response is perfect, alleviating the burden of extensive sampling and filtering to collect perfect responses. Besides, reversal also enlarges the gap between chosen and rejected responses, thereby clarifying the optimization direction and making it more robust to noise. We evaluate RPO on two multi-turn IF benchmarks, Sysbench and Multi-IF, demonstrating average improvements over the DPO baseline of 4.6 and 2.5 points (on Llama-3.1 8B), respectively. Moreover, RPO scales effectively across model sizes (8B to 70B parameters), with the 70B RPO model surpassing GPT-4o.
title Reverse Preference Optimization for Complex Instruction Following
topic Computation and Language
url https://arxiv.org/abs/2505.22172