Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Viano, Luca, Zhou, Ruida, Sun, Yifan, Namazifar, Mahdi, Cevher, Volkan, Sabach, Shoham, Ghavamzadeh, Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Directional-Clamp PPO
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
von: Barla, Adam, et al.
Veröffentlicht: (2026)
von: Barla, Adam, et al.
Veröffentlicht: (2026)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2026)
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2026)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
von: Pipano, Idan, et al.
Veröffentlicht: (2026)
von: Pipano, Idan, et al.
Veröffentlicht: (2026)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
von: Viel, Stefano, et al.
Veröffentlicht: (2025)
von: Viel, Stefano, et al.
Veröffentlicht: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
von: Viano, Luca, et al.
Veröffentlicht: (2024)
von: Viano, Luca, et al.
Veröffentlicht: (2024)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
C2-DPO: Constrained Controlled Direct Preference Optimization
von: Asadi, Kavosh, et al.
Veröffentlicht: (2025)
von: Asadi, Kavosh, et al.
Veröffentlicht: (2025)
Krylov Cubic Regularized Newton: A Subspace Second-Order Method with Dimension-Free Convergence Rate
von: Jiang, Ruichen, et al.
Veröffentlicht: (2024)
von: Jiang, Ruichen, et al.
Veröffentlicht: (2024)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
Rate optimal learning of equilibria from data
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
von: Freihaut, Till, et al.
Veröffentlicht: (2025)
Best of Both Worlds: Regret Minimization versus Minimax Play
von: Müller, Adrian, et al.
Veröffentlicht: (2025)
von: Müller, Adrian, et al.
Veröffentlicht: (2025)
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
von: Viano, Luca, et al.
Veröffentlicht: (2026)
von: Viano, Luca, et al.
Veröffentlicht: (2026)
A Proximal Operator for Inducing 2:4-Sparsity
von: Kübler, Jonas M, et al.
Veröffentlicht: (2025)
von: Kübler, Jonas M, et al.
Veröffentlicht: (2025)
FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain
von: Deb, Rohan, et al.
Veröffentlicht: (2025)
von: Deb, Rohan, et al.
Veröffentlicht: (2025)
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe
von: Thekumparampil, Kiran Koshy, et al.
Veröffentlicht: (2024)
von: Thekumparampil, Kiran Koshy, et al.
Veröffentlicht: (2024)
Optimistic Dual Averaging Unifies Modern Optimizers
von: Pethick, Thomas, et al.
Veröffentlicht: (2026)
von: Pethick, Thomas, et al.
Veröffentlicht: (2026)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
von: Liu, Fanghui, et al.
Veröffentlicht: (2024)
von: Liu, Fanghui, et al.
Veröffentlicht: (2024)
SAMPa: Sharpness-aware Minimization Parallelized
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
von: Xie, Wanyun, et al.
Veröffentlicht: (2026)
von: Xie, Wanyun, et al.
Veröffentlicht: (2026)
Learning the Target Network in Function Space
von: Asadi, Kavosh, et al.
Veröffentlicht: (2024)
von: Asadi, Kavosh, et al.
Veröffentlicht: (2024)
Improving SAM Requires Rethinking its Optimization Formulation
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
von: Xie, Wanyun, et al.
Veröffentlicht: (2025)
von: Xie, Wanyun, et al.
Veröffentlicht: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
von: Pethick, Thomas, et al.
Veröffentlicht: (2023)
von: Pethick, Thomas, et al.
Veröffentlicht: (2023)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
von: Afzal, Arshia, et al.
Veröffentlicht: (2025)
von: Afzal, Arshia, et al.
Veröffentlicht: (2025)
Inverse Q-Learning Done Right: Offline Imitation Learning in $Q^π$-Realizable MDPs
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
von: Moulin, Antoine, et al.
Veröffentlicht: (2025)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
Easy Data Unlearning Bench
von: Rinberg, Roy, et al.
Veröffentlicht: (2026)
von: Rinberg, Roy, et al.
Veröffentlicht: (2026)
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2026)
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2026)
Adversarial Training for Defense Against Label Poisoning Attacks
von: Bal, Melis Ilayda, et al.
Veröffentlicht: (2025)
von: Bal, Melis Ilayda, et al.
Veröffentlicht: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
von: Bal, Melis Ilayda, et al.
Veröffentlicht: (2025)
von: Bal, Melis Ilayda, et al.
Veröffentlicht: (2025)
Bayesian Regret Minimization in Offline Bandits
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
Bayesian policy gradient and actor-critic algorithms
von: Ghavamzadeh, Mohammad, et al.
Veröffentlicht: (2026)
von: Ghavamzadeh, Mohammad, et al.
Veröffentlicht: (2026)
Contextual Bandits with Stage-wise Constraints
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
von: Bergerault, Antoine, et al.
Veröffentlicht: (2026)
von: Bergerault, Antoine, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Directional-Clamp PPO
von: Karpel, Gilad, et al.
Veröffentlicht: (2025) -
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
von: Barla, Adam, et al.
Veröffentlicht: (2026) -
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026) -
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
von: Sheebaelhamd, Ziyad, et al.
Veröffentlicht: (2026) -
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
von: Pipano, Idan, et al.
Veröffentlicht: (2026)