Active Preference Optimization for Sample Efficient RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Das, Nirjhar, Chakraborty, Souradip, Pacchiano, Aldo, Chowdhury, Sayak Ray |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
von: Xie, Tengyang, et al.
Veröffentlicht: (2024)
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
von: Misra, Dipendra, et al.
Veröffentlicht: (2026)
von: Misra, Dipendra, et al.
Veröffentlicht: (2026)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Post-training Large Language Models for Diverse High-Quality Responses
von: Chen, Yilei, et al.
Veröffentlicht: (2025)
von: Chen, Yilei, et al.
Veröffentlicht: (2025)
When Less is Enough: Efficient Inference via Collaborative Reasoning
von: Chen, Yilei, et al.
Veröffentlicht: (2026)
von: Chen, Yilei, et al.
Veröffentlicht: (2026)
Provable Interactive Learning with Hindsight Instruction Feedback
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
Adaptive Margin RLHF via Preference over Preferences
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
von: Dang, John, et al.
Veröffentlicht: (2024)
von: Dang, John, et al.
Veröffentlicht: (2024)
A Theoretical Framework for Partially Observed Reward-States in RLHF
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
Dataset Reset Policy Optimization for RLHF
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
von: Du, Yihan, et al.
Veröffentlicht: (2024)
von: Du, Yihan, et al.
Veröffentlicht: (2024)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
Second Order Bounds for Contextual Bandits with Function Approximation
von: Pacchiano, Aldo
Veröffentlicht: (2024)
von: Pacchiano, Aldo
Veröffentlicht: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
von: Ding, Mucong, et al.
Veröffentlicht: (2024)
DPO Meets PPO: Reinforced Token Optimization for RLHF
von: Zhong, Han, et al.
Veröffentlicht: (2024)
von: Zhong, Han, et al.
Veröffentlicht: (2024)
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
von: Melikidze, Davit, et al.
Veröffentlicht: (2026)
von: Melikidze, Davit, et al.
Veröffentlicht: (2026)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
von: Patel, Bhrij, et al.
Veröffentlicht: (2024)
Time-Constrained Recommendations: Reinforcement Learning Strategies for E-Commerce
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2025)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2025)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
von: Hu, Jian, et al.
Veröffentlicht: (2024)
von: Hu, Jian, et al.
Veröffentlicht: (2024)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2024)
Agentic Critical Training
von: Liu, Weize, et al.
Veröffentlicht: (2026)
von: Liu, Weize, et al.
Veröffentlicht: (2026)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
RLHF and IIA: Perverse Incentives
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
von: Belakaria, Syrine, et al.
Veröffentlicht: (2025)
von: Belakaria, Syrine, et al.
Veröffentlicht: (2025)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Lin, Xiaofeng, et al.
Veröffentlicht: (2026)
On the Role of Preference Variance in Preference Optimization
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
The Perfect Blend: Redefining RLHF with Mixture of Judges
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
von: Xu, Tengyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024) -
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
von: Xie, Tengyang, et al.
Veröffentlicht: (2024) -
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
von: Chakrabarty, Sayak, et al.
Veröffentlicht: (2024) -
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
von: Misra, Dipendra, et al.
Veröffentlicht: (2026) -
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)