Continuous-Utility Direct Preference Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mohsin, Muhammad Ahmed, Umer, Muhammad, Bilal, Ahsan, He, Zihao, Rafique, Muhammad Usman, Aali, Asad, Jamshed, Muhammad Ali, Cioffi, John M., Fox, Emily
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910160154263552
author Mohsin, Muhammad Ahmed
Umer, Muhammad
Bilal, Ahsan
He, Zihao
Rafique, Muhammad Usman
Aali, Asad
Jamshed, Muhammad Ali
Cioffi, John M.
Fox, Emily
author_facet Mohsin, Muhammad Ahmed
Umer, Muhammad
Bilal, Ahsan
He, Zihao
Rafique, Muhammad Usman
Aali, Asad
Jamshed, Muhammad Ali
Cioffi, John M.
Fox, Emily
contents Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasoning quality. We introduce Continuous Utility Direct Preference Optimization (CU-DPO), a framework that aligns models to a portfolio of prompt-based cognitive strategies by replacing binary labels with continuous scores that capture fine-grained reasoning quality. We prove that learning with K strategies yields a Theta(K log K) improvement in sample complexity over binary preferences, and that DPO converges to the entropy-regularized utility-maximizing policy. To exploit this signal, we propose a two-stage training pipeline: (i) strategy selection, which optimizes the model to choose the best strategy for a given problem via best-vs-all comparisons, and (ii) execution refinement, which trains the model to correctly execute the selected strategy using margin-stratified pairs. On mathematical reasoning benchmarks, CU-DPO improves strategy selection accuracy from 35-46 percent to 68-78 percent across seven base models, yielding consistent downstream reasoning gains of up to 6.6 points on in-distribution datasets with effective transfer to out-of-distribution tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00931
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Continuous-Utility Direct Preference Optimization
Mohsin, Muhammad Ahmed
Umer, Muhammad
Bilal, Ahsan
He, Zihao
Rafique, Muhammad Usman
Aali, Asad
Jamshed, Muhammad Ali
Cioffi, John M.
Fox, Emily
Machine Learning
Artificial Intelligence
Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasoning quality. We introduce Continuous Utility Direct Preference Optimization (CU-DPO), a framework that aligns models to a portfolio of prompt-based cognitive strategies by replacing binary labels with continuous scores that capture fine-grained reasoning quality. We prove that learning with K strategies yields a Theta(K log K) improvement in sample complexity over binary preferences, and that DPO converges to the entropy-regularized utility-maximizing policy. To exploit this signal, we propose a two-stage training pipeline: (i) strategy selection, which optimizes the model to choose the best strategy for a given problem via best-vs-all comparisons, and (ii) execution refinement, which trains the model to correctly execute the selected strategy using margin-stratified pairs. On mathematical reasoning benchmarks, CU-DPO improves strategy selection accuracy from 35-46 percent to 68-78 percent across seven base models, yielding consistent downstream reasoning gains of up to 6.6 points on in-distribution datasets with effective transfer to out-of-distribution tasks.
title Continuous-Utility Direct Preference Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.00931