Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cen, Shicong, Mei, Jincheng, Goshvadi, Katayoon, Dai, Hanjun, Yang, Tong, Yang, Sherry, Schuurmans, Dale, Chi, Yuejie, Dai, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Expectations: Learning with Stochastic Dominance Made Practical
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
von: Yang, Tong, et al.
Veröffentlicht: (2024)
von: Yang, Tong, et al.
Veröffentlicht: (2024)
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
von: Yang, Tong, et al.
Veröffentlicht: (2025)
von: Yang, Tong, et al.
Veröffentlicht: (2025)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
von: Yang, Tong, et al.
Veröffentlicht: (2025)
von: Yang, Tong, et al.
Veröffentlicht: (2025)
Autoregressive Large Language Models are Computationally Universal
von: Schuurmans, Dale, et al.
Veröffentlicht: (2024)
von: Schuurmans, Dale, et al.
Veröffentlicht: (2024)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
von: Yang, Tong, et al.
Veröffentlicht: (2023)
von: Yang, Tong, et al.
Veröffentlicht: (2023)
Judging with Confidence: Calibrating Autoraters to Preference Distributions
von: Li, Zhuohang, et al.
Veröffentlicht: (2025)
von: Li, Zhuohang, et al.
Veröffentlicht: (2025)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Stochastic Gradient Succeeds for Bandits
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning
von: Dai, Bo, et al.
Veröffentlicht: (2026)
von: Dai, Bo, et al.
Veröffentlicht: (2026)
UQE: A Query Engine for Unstructured Databases
von: Dai, Hanjun, et al.
Veröffentlicht: (2024)
von: Dai, Hanjun, et al.
Veröffentlicht: (2024)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Online Bandit Learning with Offline Preference Data for Improved RLHF
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024)
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024)
Large Language Models can Learn Rules
von: Zhu, Zhaocheng, et al.
Veröffentlicht: (2023)
von: Zhu, Zhaocheng, et al.
Veröffentlicht: (2023)
Preference Optimization for Molecule Synthesis with Conditional Residual Energy-based Models
von: Liu, Songtao, et al.
Veröffentlicht: (2024)
von: Liu, Songtao, et al.
Veröffentlicht: (2024)
Offline Constrained RLHF with Multiple Preference Oracles
von: Latham, Brenden, et al.
Veröffentlicht: (2026)
von: Latham, Brenden, et al.
Veröffentlicht: (2026)
Spectral Representation-based Reinforcement Learning
von: Gao, Chenxiao, et al.
Veröffentlicht: (2025)
von: Gao, Chenxiao, et al.
Veröffentlicht: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
von: Zhang, Hongming, et al.
Veröffentlicht: (2023)
von: Zhang, Hongming, et al.
Veröffentlicht: (2023)
Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity
von: Shi, Laixi, et al.
Veröffentlicht: (2022)
von: Shi, Laixi, et al.
Veröffentlicht: (2022)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
Soft Preference Optimization: Aligning Language Models to Expert Distributions
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2024)
von: Sharifnassab, Arsalan, et al.
Veröffentlicht: (2024)
Generalized Preference Optimization: A Unified Approach to Offline Alignment
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
von: Tang, Yunhao, et al.
Veröffentlicht: (2024)
Unified Dual Semidefinite Programming Framework for Incentive-Compatible Mechanism Design
von: Zhang, Jincheng
Veröffentlicht: (2026)
von: Zhang, Jincheng
Veröffentlicht: (2026)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
Sample Complexity of Offline Distributionally Robust Linear Markov Decision Processes
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
von: Che, Fengdi, et al.
Veröffentlicht: (2024)
von: Che, Fengdi, et al.
Veröffentlicht: (2024)
Regularized Online RLHF with Generalized Bilinear Preferences
von: Lee, Junghyun, et al.
Veröffentlicht: (2026)
von: Lee, Junghyun, et al.
Veröffentlicht: (2026)
Spectral Souping: A Unified Framework for Online Preference Alignment
von: Chow, Yinlam, et al.
Veröffentlicht: (2026)
von: Chow, Yinlam, et al.
Veröffentlicht: (2026)
Offline and Online KL-Regularized RLHF under Differential Privacy
von: Wu, Yulian, et al.
Veröffentlicht: (2025)
von: Wu, Yulian, et al.
Veröffentlicht: (2025)
RLHF and IIA: Perverse Incentives
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
Diffusion Controller: Framework, Algorithms and Parameterization
von: Yang, Tong, et al.
Veröffentlicht: (2026)
von: Yang, Tong, et al.
Veröffentlicht: (2026)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
von: Dai, Juntao, et al.
Veröffentlicht: (2025)
von: Dai, Juntao, et al.
Veröffentlicht: (2025)
Representation Learning via Non-Contrastive Mutual Information
von: Guo, Zhaohan Daniel, et al.
Veröffentlicht: (2025)
von: Guo, Zhaohan Daniel, et al.
Veröffentlicht: (2025)
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2025)
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2025)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
von: Mukherjee, Arpan, et al.
Veröffentlicht: (2025)
von: Mukherjee, Arpan, et al.
Veröffentlicht: (2025)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Federated Offline Reinforcement Learning: Collaborative Single-Policy Coverage Suffices
von: Woo, Jiin, et al.
Veröffentlicht: (2024)
von: Woo, Jiin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Expectations: Learning with Stochastic Dominance Made Practical
von: Cen, Shicong, et al.
Veröffentlicht: (2024) -
Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
von: Yang, Tong, et al.
Veröffentlicht: (2024) -
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL
von: Yang, Tong, et al.
Veröffentlicht: (2025) -
Exploring and Benchmarking the Planning Capabilities of Large Language Models
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024) -
Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov Games
von: Yang, Tong, et al.
Veröffentlicht: (2025)