Reverse Preference Optimization for Complex Instruction Following
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Xiang, Lin, Ting-En, Fang, Feiteng, Wu, Yuchuan, Li, Hangyu, Qu, Yuzhong, Huang, Fei, Li, Yongbin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
by: Ye, Xinge, et al.
Published: (2025)
by: Ye, Xinge, et al.
Published: (2025)
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
by: Zhang, Xinghua, et al.
Published: (2024)
by: Zhang, Xinghua, et al.
Published: (2024)
Improving Factual Consistency of News Summarization by Contrastive Preference Optimization
by: Feng, Huawen, et al.
Published: (2023)
by: Feng, Huawen, et al.
Published: (2023)
EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
by: Zou, Tao, et al.
Published: (2025)
by: Zou, Tao, et al.
Published: (2025)
SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
by: Si, Shuzheng, et al.
Published: (2023)
by: Si, Shuzheng, et al.
Published: (2023)
A Survey on Self-Evolution of Large Language Models
by: Tao, Zhengwei, et al.
Published: (2024)
by: Tao, Zhengwei, et al.
Published: (2024)
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
by: Zhang, Haonan, et al.
Published: (2025)
by: Zhang, Haonan, et al.
Published: (2025)
SDPO: Segment-Level Direct Preference Optimization for Social Agents
by: Kong, Aobo, et al.
Published: (2025)
by: Kong, Aobo, et al.
Published: (2025)
Act-Adaptive Margin: Dynamically Calibrating Reward Models for Subjective Ambiguity
by: Fang, Feiteng, et al.
Published: (2025)
by: Fang, Feiteng, et al.
Published: (2025)
Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use
by: Chen, Yuhan, et al.
Published: (2023)
by: Chen, Yuhan, et al.
Published: (2023)
MOA: Multi-Objective Alignment for Role-Playing Agents
by: Liao, Chonghua, et al.
Published: (2025)
by: Liao, Chonghua, et al.
Published: (2025)
P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling
by: Zhang, Pinyi, et al.
Published: (2026)
by: Zhang, Pinyi, et al.
Published: (2026)
Preference Ranking Optimization for Human Alignment
by: Song, Feifan, et al.
Published: (2023)
by: Song, Feifan, et al.
Published: (2023)
Agentic Reinforcement Learning with Implicit Step Rewards
by: Liu, Xiaoqian, et al.
Published: (2025)
by: Liu, Xiaoqian, et al.
Published: (2025)
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning
by: Liu, Xiaoqian, et al.
Published: (2025)
by: Liu, Xiaoqian, et al.
Published: (2025)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
by: Xiao, Ruixuan, et al.
Published: (2024)
by: Xiao, Ruixuan, et al.
Published: (2024)
Supervised Optimism Correction: Be Confident When LLMs Are Sure
by: Zhang, Junjie, et al.
Published: (2025)
by: Zhang, Junjie, et al.
Published: (2025)
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
by: Zhou, Qinhao, et al.
Published: (2024)
by: Zhou, Qinhao, et al.
Published: (2024)
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Benchmarking Complex Instruction-Following with Multiple Constraints Composition
by: Wen, Bosi, et al.
Published: (2024)
by: Wen, Bosi, et al.
Published: (2024)
Debate Helps Weak-to-Strong Generalization
by: Lang, Hao, et al.
Published: (2025)
by: Lang, Hao, et al.
Published: (2025)
Selective Weak-to-Strong Generalization
by: Lang, Hao, et al.
Published: (2025)
by: Lang, Hao, et al.
Published: (2025)
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
by: Zhang, Junjie, et al.
Published: (2025)
by: Zhang, Junjie, et al.
Published: (2025)
TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence
by: Hou, Guiyang, et al.
Published: (2025)
by: Hou, Guiyang, et al.
Published: (2025)
Enhancing Complex Question Answering over Knowledge Graphs through Evidence Pattern Retrieval
by: Ding, Wentao, et al.
Published: (2024)
by: Ding, Wentao, et al.
Published: (2024)
Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language Models
by: Chen, Changyu, et al.
Published: (2024)
by: Chen, Changyu, et al.
Published: (2024)
Fine-Tuning Language Models with Reward Learning on Policy
by: Lang, Hao, et al.
Published: (2024)
by: Lang, Hao, et al.
Published: (2024)
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
by: Luo, Run, et al.
Published: (2024)
by: Luo, Run, et al.
Published: (2024)
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
TARGA: Targeted Synthetic Data Generation for Practical Reasoning over Structured Data
by: Huang, Xiang, et al.
Published: (2024)
by: Huang, Xiang, et al.
Published: (2024)
M3PO: Multimodal-Model-Guided Preference Optimization for Visual Instruction Following
by: Gao, Ruirui, et al.
Published: (2025)
by: Gao, Ruirui, et al.
Published: (2025)
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment
by: Zhao, Yingxiu, et al.
Published: (2023)
by: Zhao, Yingxiu, et al.
Published: (2023)
Extend Model Merging from Fine-Tuned to Pre-Trained Large Language Models via Weight Disentanglement
by: Yu, Le, et al.
Published: (2024)
by: Yu, Le, et al.
Published: (2024)
DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever
by: Yin, Zhichao, et al.
Published: (2024)
by: Yin, Zhichao, et al.
Published: (2024)
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
by: Huang, Hui, et al.
Published: (2025)
by: Huang, Hui, et al.
Published: (2025)
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
QueryAgent: A Reliable and Efficient Reasoning Framework with Environmental Feedback-based Self-Correction
by: Huang, Xiang, et al.
Published: (2024)
by: Huang, Xiang, et al.
Published: (2024)
DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking
by: Li, Zhuoqun, et al.
Published: (2025)
by: Li, Zhuoqun, et al.
Published: (2025)
BPO: Revisiting Preference Modeling in Direct Preference Optimization
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
Similar Items
-
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
by: Ye, Xinge, et al.
Published: (2025) -
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
by: Zhang, Xinghua, et al.
Published: (2024) -
Improving Factual Consistency of News Summarization by Contrastive Preference Optimization
by: Feng, Huawen, et al.
Published: (2023) -
EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
by: Zou, Tao, et al.
Published: (2025) -
SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
by: Si, Shuzheng, et al.
Published: (2023)