Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Tengyang, Foster, Dylan J., Krishnamurthy, Akshay, Rosset, Corby, Awadallah, Ahmed, Rakhlin, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
von: Mitra, Arindam, et al.
Veröffentlicht: (2024)
von: Mitra, Arindam, et al.
Veröffentlicht: (2024)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
von: Das, Nirjhar, et al.
Veröffentlicht: (2024)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
von: Jia, Zeyu, et al.
Veröffentlicht: (2025)
von: Jia, Zeyu, et al.
Veröffentlicht: (2025)
Fara-7B: An Efficient Agentic Model for Computer Use
von: Awadallah, Ahmed, et al.
Veröffentlicht: (2025)
von: Awadallah, Ahmed, et al.
Veröffentlicht: (2025)
The Art of Building Verifiers for Computer Use Agents
von: Rosset, Corby, et al.
Veröffentlicht: (2026)
von: Rosset, Corby, et al.
Veröffentlicht: (2026)
Can large language models explore in-context?
von: Krishnamurthy, Akshay, et al.
Veröffentlicht: (2024)
von: Krishnamurthy, Akshay, et al.
Veröffentlicht: (2024)
WPO: Enhancing RLHF with Weighted Preference Optimization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
Automatic Pair Construction for Contrastive Post-training
von: Xu, Canwen, et al.
Veröffentlicht: (2023)
von: Xu, Canwen, et al.
Veröffentlicht: (2023)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
von: Chen, Fan, et al.
Veröffentlicht: (2025)
von: Chen, Fan, et al.
Veröffentlicht: (2025)
General Exploratory Bonus for Optimistic Exploration in RLHF
von: Li, Wendi, et al.
Veröffentlicht: (2025)
von: Li, Wendi, et al.
Veröffentlicht: (2025)
The Power of Resets in Online Reinforcement Learning
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2024)
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2024)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
von: Yuan, Yurun, et al.
Veröffentlicht: (2026)
Reinforce LLM Reasoning through Multi-Agent Reflection
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)
Adaptive Margin RLHF via Preference over Preferences
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
von: Li, Aaron J., et al.
Veröffentlicht: (2024)
von: Li, Aaron J., et al.
Veröffentlicht: (2024)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
von: Dang, John, et al.
Veröffentlicht: (2024)
von: Dang, John, et al.
Veröffentlicht: (2024)
Self-Improvement in Language Models: The Sharpening Mechanism
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents
von: Pahuja, Vardaan, et al.
Veröffentlicht: (2025)
von: Pahuja, Vardaan, et al.
Veröffentlicht: (2025)
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
von: Yang, Wen, et al.
Veröffentlicht: (2025)
von: Yang, Wen, et al.
Veröffentlicht: (2025)
AgentInstruct: Toward Generative Teaching with Agentic Flows
von: Mitra, Arindam, et al.
Veröffentlicht: (2024)
von: Mitra, Arindam, et al.
Veröffentlicht: (2024)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
Prototypical Reward Network for Data-Efficient RLHF
von: Zhang, Jinghan, et al.
Veröffentlicht: (2024)
von: Zhang, Jinghan, et al.
Veröffentlicht: (2024)
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
von: Golowich, Noah, et al.
Veröffentlicht: (2026)
von: Golowich, Noah, et al.
Veröffentlicht: (2026)
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
von: Tuyls, Jens, et al.
Veröffentlicht: (2025)
von: Tuyls, Jens, et al.
Veröffentlicht: (2025)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
MaxMin-RLHF: Alignment with Diverse Human Preferences
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
The Coverage Principle: How Pre-Training Enables Post-Training
von: Chen, Fan, et al.
Veröffentlicht: (2025)
von: Chen, Fan, et al.
Veröffentlicht: (2025)
From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models
von: Raheja, Tarun, et al.
Veröffentlicht: (2026)
von: Raheja, Tarun, et al.
Veröffentlicht: (2026)
Dataset Reset Policy Optimization for RLHF
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
von: Du, Yihan, et al.
Veröffentlicht: (2024)
von: Du, Yihan, et al.
Veröffentlicht: (2024)
Preference Packing: Efficient Preference Optimization for Large Language Models
von: Cho, Jaekyung
Veröffentlicht: (2026)
von: Cho, Jaekyung
Veröffentlicht: (2026)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
von: Singh, Joykirat, et al.
Veröffentlicht: (2025)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024) -
Orca-Math: Unlocking the potential of SLMs in Grade School Math
von: Mitra, Arindam, et al.
Veröffentlicht: (2024) -
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
von: Huang, Audrey, et al.
Veröffentlicht: (2024) -
Active Preference Optimization for Sample Efficient RLHF
von: Das, Nirjhar, et al.
Veröffentlicht: (2024) -
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
von: Yuan, Yurun, et al.
Veröffentlicht: (2025)