Continuous-Utility Direct Preference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910160154263552 |
|---|---|
| author | Mohsin, Muhammad Ahmed Umer, Muhammad Bilal, Ahsan He, Zihao Rafique, Muhammad Usman Aali, Asad Jamshed, Muhammad Ali Cioffi, John M. Fox, Emily |
| author_facet | Mohsin, Muhammad Ahmed Umer, Muhammad Bilal, Ahsan He, Zihao Rafique, Muhammad Usman Aali, Asad Jamshed, Muhammad Ali Cioffi, John M. Fox, Emily |
| contents | Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasoning quality. We introduce Continuous Utility Direct Preference Optimization (CU-DPO), a framework that aligns models to a portfolio of prompt-based cognitive strategies by replacing binary labels with continuous scores that capture fine-grained reasoning quality. We prove that learning with K strategies yields a Theta(K log K) improvement in sample complexity over binary preferences, and that DPO converges to the entropy-regularized utility-maximizing policy. To exploit this signal, we propose a two-stage training pipeline: (i) strategy selection, which optimizes the model to choose the best strategy for a given problem via best-vs-all comparisons, and (ii) execution refinement, which trains the model to correctly execute the selected strategy using margin-stratified pairs. On mathematical reasoning benchmarks, CU-DPO improves strategy selection accuracy from 35-46 percent to 68-78 percent across seven base models, yielding consistent downstream reasoning gains of up to 6.6 points on in-distribution datasets with effective transfer to out-of-distribution tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_00931 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Continuous-Utility Direct Preference Optimization Mohsin, Muhammad Ahmed Umer, Muhammad Bilal, Ahsan He, Zihao Rafique, Muhammad Usman Aali, Asad Jamshed, Muhammad Ali Cioffi, John M. Fox, Emily Machine Learning Artificial Intelligence Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasoning quality. We introduce Continuous Utility Direct Preference Optimization (CU-DPO), a framework that aligns models to a portfolio of prompt-based cognitive strategies by replacing binary labels with continuous scores that capture fine-grained reasoning quality. We prove that learning with K strategies yields a Theta(K log K) improvement in sample complexity over binary preferences, and that DPO converges to the entropy-regularized utility-maximizing policy. To exploit this signal, we propose a two-stage training pipeline: (i) strategy selection, which optimizes the model to choose the best strategy for a given problem via best-vs-all comparisons, and (ii) execution refinement, which trains the model to correctly execute the selected strategy using margin-stratified pairs. On mathematical reasoning benchmarks, CU-DPO improves strategy selection accuracy from 35-46 percent to 68-78 percent across seven base models, yielding consistent downstream reasoning gains of up to 6.6 points on in-distribution datasets with effective transfer to out-of-distribution tasks. |
| title | Continuous-Utility Direct Preference Optimization |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2602.00931 |