Value-Free Policy Optimization via Reward Partitioning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Faye, Bilal, Azzag, Hanane, Lebbah, Mustapha
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917553949900800
author Faye, Bilal
Azzag, Hanane
Lebbah, Mustapha
author_facet Faye, Bilal
Azzag, Hanane
Lebbah, Mustapha
contents Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative to pairwise preference learning by directly leveraging scalar feedback. Existing approaches such as Direct Reward Optimization (DRO) have demonstrated promising results but rely on value function estimation, introducing additional variance, optimization complexity, and sensitivity to off-policy data. We introduce Reward Partition Optimization (RPO), a simple and scalable reward-driven objective that eliminates the need for value function learning. RPO normalizes rewards through a partition-based formulation estimated directly from prompt-level reward distributions, yielding a stable supervised optimization objective without auxiliary models or reinforcement learning loops. We evaluate RPO across multiple encoder-decoder and decoder-only language models using automatic metrics, LLM-as-a-judge evaluations, and optimization stability analyses. Experimental results show that RPO consistently outperforms strong baselines, including SFT, KTO, and DRO, while producing more aligned, diverse, and less toxic generations.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13702
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Value-Free Policy Optimization via Reward Partitioning
Faye, Bilal
Azzag, Hanane
Lebbah, Mustapha
Machine Learning
Artificial Intelligence
Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative to pairwise preference learning by directly leveraging scalar feedback. Existing approaches such as Direct Reward Optimization (DRO) have demonstrated promising results but rely on value function estimation, introducing additional variance, optimization complexity, and sensitivity to off-policy data. We introduce Reward Partition Optimization (RPO), a simple and scalable reward-driven objective that eliminates the need for value function learning. RPO normalizes rewards through a partition-based formulation estimated directly from prompt-level reward distributions, yielding a stable supervised optimization objective without auxiliary models or reinforcement learning loops. We evaluate RPO across multiple encoder-decoder and decoder-only language models using automatic metrics, LLM-as-a-judge evaluations, and optimization stability analyses. Experimental results show that RPO consistently outperforms strong baselines, including SFT, KTO, and DRO, while producing more aligned, diverse, and less toxic generations.
title Value-Free Policy Optimization via Reward Partitioning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.13702