RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Chanwoo, Liu, Mingyang, Kong, Dingwen, Zhang, Kaiqing, Ozdaglar, Asuman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
von: Park, Chanwoo, et al.
Veröffentlicht: (2024)
von: Park, Chanwoo, et al.
Veröffentlicht: (2024)
Multi-Player Zero-Sum Markov Games with Networked Separable Interactions
von: Park, Chanwoo, et al.
Veröffentlicht: (2023)
von: Park, Chanwoo, et al.
Veröffentlicht: (2023)
Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
von: Kim, Kihyun, et al.
Veröffentlicht: (2025)
LiteEFG: An Efficient Python Library for Solving Extensive-form Games
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)
A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)
Differentially Private Equilibrium Finding in Polymatrix Games
von: Liu, Mingyang, et al.
Veröffentlicht: (2025)
von: Liu, Mingyang, et al.
Veröffentlicht: (2025)
A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback
von: Kim, Kihyun, et al.
Veröffentlicht: (2024)
von: Kim, Kihyun, et al.
Veröffentlicht: (2024)
Computing Equilibrium beyond Unilateral Deviation
von: Liu, Mingyang, et al.
Veröffentlicht: (2026)
von: Liu, Mingyang, et al.
Veröffentlicht: (2026)
The Power of Regularization in Solving Extensive-Form Games
von: Liu, Mingyang, et al.
Veröffentlicht: (2022)
von: Liu, Mingyang, et al.
Veröffentlicht: (2022)
MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
Online Learning and Equilibrium Computation with Ranking Feedback
von: Liu, Mingyang, et al.
Veröffentlicht: (2026)
von: Liu, Mingyang, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning via Linear-Programming with Error-Bound Induced Constraints
von: Ozdaglar, Asuman, et al.
Veröffentlicht: (2022)
von: Ozdaglar, Asuman, et al.
Veröffentlicht: (2022)
T-POP: Test-Time Personalization with Online Preference Feedback
von: Qu, Zikun, et al.
Veröffentlicht: (2025)
von: Qu, Zikun, et al.
Veröffentlicht: (2025)
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
von: Xiong, Wei, et al.
Veröffentlicht: (2023)
von: Xiong, Wei, et al.
Veröffentlicht: (2023)
UFT: Unifying Supervised and Reinforcement Fine-Tuning
von: Liu, Mingyang, et al.
Veröffentlicht: (2025)
von: Liu, Mingyang, et al.
Veröffentlicht: (2025)
Adaptive Margin RLHF via Preference over Preferences
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
von: Cho, Taehyun, et al.
Veröffentlicht: (2025)
von: Cho, Taehyun, et al.
Veröffentlicht: (2025)
A Shared Low-Rank Adaptation Approach to Personalized RLHF
von: Liu, Renpu, et al.
Veröffentlicht: (2025)
von: Liu, Renpu, et al.
Veröffentlicht: (2025)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Democratic Preference Alignment via Sortition-Weighted RLHF
von: Sana, Suvadip, et al.
Veröffentlicht: (2026)
von: Sana, Suvadip, et al.
Veröffentlicht: (2026)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
von: Belakaria, Syrine, et al.
Veröffentlicht: (2025)
von: Belakaria, Syrine, et al.
Veröffentlicht: (2025)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
Last-Iterate Convergence of Payoff-Based Independent Learning in Zero-Sum Stochastic Games
von: Chen, Zaiwei, et al.
Veröffentlicht: (2024)
von: Chen, Zaiwei, et al.
Veröffentlicht: (2024)
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
von: Ji, Jiaming, et al.
Veröffentlicht: (2025)
von: Ji, Jiaming, et al.
Veröffentlicht: (2025)
Avoiding $\mathbf{exp(R_{max})}$ scaling in RLHF through Preference-based Exploration
von: Chen, Mingyu, et al.
Veröffentlicht: (2025)
von: Chen, Mingyu, et al.
Veröffentlicht: (2025)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
von: Kim, Gihoon, et al.
Veröffentlicht: (2026)
von: Kim, Gihoon, et al.
Veröffentlicht: (2026)
APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2026)
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2026)
Semi-Supervised Preference Optimization with Limited Feedback
von: Lee, Seonggyun, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyun, et al.
Veröffentlicht: (2025)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
Learning Personalized Agents from Human Feedback
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
Influencing Humans to Conform to Preference Models for RLHF
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
von: Zhang, Rongtao, et al.
Veröffentlicht: (2026)
von: Zhang, Rongtao, et al.
Veröffentlicht: (2026)
ROCM: RLHF on consistency models
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2025)
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2025)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
LLM Misalignment via Adversarial RLHF Platforms
von: Entezami, Erfan, et al.
Veröffentlicht: (2025)
von: Entezami, Erfan, et al.
Veröffentlicht: (2025)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
von: Park, Chanwoo, et al.
Veröffentlicht: (2024) -
Multi-Player Zero-Sum Markov Games with Networked Separable Interactions
von: Park, Chanwoo, et al.
Veröffentlicht: (2023) -
Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach
von: Park, Chanwoo, et al.
Veröffentlicht: (2025) -
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
von: Kim, Kihyun, et al.
Veröffentlicht: (2025) -
LiteEFG: An Efficient Python Library for Solving Extensive-form Games
von: Liu, Mingyang, et al.
Veröffentlicht: (2024)