Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nika, Andi, Mandal, Debmalya, Kamalaruban, Parameswaran, Tzannetos, Georgios, Radanović, Goran, Singla, Adish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Policy Teaching via Data Poisoning in Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2025)
von: Nika, Andi, et al.
Veröffentlicht: (2025)
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
von: Nika, Andi, et al.
Veröffentlicht: (2026)
von: Nika, Andi, et al.
Veröffentlicht: (2026)
Corruption Robust Offline Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024)
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
von: Nika, Andi, et al.
Veröffentlicht: (2024)
von: Nika, Andi, et al.
Veröffentlicht: (2024)
Learning Embeddings for Sequential Tasks Using Population of Agents
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023)
von: Mahajan, Mridul, et al.
Veröffentlicht: (2023)
Informativeness of Reward Functions in Reinforcement Learning
von: Devidze, Rati, et al.
Veröffentlicht: (2024)
von: Devidze, Rati, et al.
Veröffentlicht: (2024)
Sparse Offline Reinforcement Learning with Corruption Robustness
von: Tran, Nam Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Nam Phuong, et al.
Veröffentlicht: (2025)
Performative Reinforcement Learning with Linear Markov Decision Process
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024)
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
Distributionally Robust Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2025)
On Corruption-Robustness in Performative Reinforcement Learning
von: Pollatos, Vasilis, et al.
Veröffentlicht: (2025)
von: Pollatos, Vasilis, et al.
Veröffentlicht: (2025)
Inference-Time Personalized Alignment with a Few User Preference Queries
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2025)
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026)
von: Nöther, Jonathan, et al.
Veröffentlicht: (2026)
Neural Task Synthesis for Visual Programming
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2023)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2023)
Performative Reinforcement Learning in Gradually Shifting Environments
von: Rank, Ben, et al.
Veröffentlicht: (2024)
von: Rank, Ben, et al.
Veröffentlicht: (2024)
Stochastic Principal-Agent Problems: Efficient Computation and Learning
von: Gan, Jiarui, et al.
Veröffentlicht: (2023)
von: Gan, Jiarui, et al.
Veröffentlicht: (2023)
Independent Learning in Performative Markov Potential Games
von: Sahitaj, Rilind, et al.
Veröffentlicht: (2025)
von: Sahitaj, Rilind, et al.
Veröffentlicht: (2025)
Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks?
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
von: Sasnauskas, Paulius, et al.
Veröffentlicht: (2025)
Reward Design for Justifiable Sequential Decision-Making
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
von: Sukovic, Aleksa, et al.
Veröffentlicht: (2024)
Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human Input
von: Peng, Andi, et al.
Veröffentlicht: (2024)
von: Peng, Andi, et al.
Veröffentlicht: (2024)
Formal Models of Active Learning from Contrastive Examples
von: Mansouri, Farnam, et al.
Veröffentlicht: (2025)
von: Mansouri, Farnam, et al.
Veröffentlicht: (2025)
Strategyproof Reinforcement Learning from Human Feedback
von: Buening, Thomas Kleine, et al.
Veröffentlicht: (2025)
von: Buening, Thomas Kleine, et al.
Veröffentlicht: (2025)
Learning Personalized Decision Support Policies
von: Bhatt, Umang, et al.
Veröffentlicht: (2023)
von: Bhatt, Umang, et al.
Veröffentlicht: (2023)
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024)
von: Kotalwar, Nachiket, et al.
Veröffentlicht: (2024)
Adversarially Robust Decision Transformer
von: Tang, Xiaohang, et al.
Veröffentlicht: (2024)
von: Tang, Xiaohang, et al.
Veröffentlicht: (2024)
Agent-Specific Effects: A Causal Effect Propagation Analysis in Multi-Agent MDPs
von: Triantafyllou, Stelios, et al.
Veröffentlicht: (2023)
von: Triantafyllou, Stelios, et al.
Veröffentlicht: (2023)
Beyond Grids: Multi-objective Bayesian Optimization With Adaptive Discretization
von: Nika, Andi, et al.
Veröffentlicht: (2020)
von: Nika, Andi, et al.
Veröffentlicht: (2020)
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
von: Gölz, Paul, et al.
Veröffentlicht: (2025)
von: Gölz, Paul, et al.
Veröffentlicht: (2025)
Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
von: Radmehr, Bahar, et al.
Veröffentlicht: (2024)
von: Radmehr, Bahar, et al.
Veröffentlicht: (2024)
Active Learning for Direct Preference Optimization
von: Kveton, Branislav, et al.
Veröffentlicht: (2025)
von: Kveton, Branislav, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Durable Algorithmic Recourse
von: Ceccon, Marina, et al.
Veröffentlicht: (2025)
von: Ceccon, Marina, et al.
Veröffentlicht: (2025)
Learning Half-Spaces from Perturbed Contrastive Examples
von: Ravari, Aryan Alavi Razavi, et al.
Veröffentlicht: (2026)
von: Ravari, Aryan Alavi Razavi, et al.
Veröffentlicht: (2026)
Evaluating Fairness in Transaction Fraud Models: Fairness Metrics, Bias Audits, and Challenges
von: Kamalaruban, Parameswaran, et al.
Veröffentlicht: (2024)
von: Kamalaruban, Parameswaran, et al.
Veröffentlicht: (2024)
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
von: Higuchi, Rei, et al.
Veröffentlicht: (2026)
von: Higuchi, Rei, et al.
Veröffentlicht: (2026)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
von: Lin, Yong, et al.
Veröffentlicht: (2024)
von: Lin, Yong, et al.
Veröffentlicht: (2024)
Contextual Combinatorial Bandits with Changing Action Sets via Gaussian Processes
von: Nika, Andi, et al.
Veröffentlicht: (2021)
von: Nika, Andi, et al.
Veröffentlicht: (2021)
Emergent Bias and Fairness in Multi-Agent Decision Systems
von: Madigan, Maeve, et al.
Veröffentlicht: (2025)
von: Madigan, Maeve, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Policy Teaching via Data Poisoning in Learning from Human Preferences
von: Nika, Andi, et al.
Veröffentlicht: (2025) -
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
von: Nika, Andi, et al.
Veröffentlicht: (2026) -
Corruption Robust Offline Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2024) -
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2024) -
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
von: Tzannetos, Georgios, et al.
Veröffentlicht: (2025)