Similar Items
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
by: Xu, Wenyuan, et al.
Published: (2025)
by: Xu, Wenyuan, et al.
Published: (2025)
Dataset Reset Policy Optimization for RLHF
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
Understanding Adversarial Imitation Learning in Small Sample Regime: A Stage-coupled Analysis
by: Xu, Tian, et al.
Published: (2022)
by: Xu, Tian, et al.
Published: (2022)
Efficient Federated RLHF via Zeroth-Order Policy Optimization
by: Wang, Deyi, et al.
Published: (2026)
by: Wang, Deyi, et al.
Published: (2026)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)
by: Du, Yihan, et al.
Published: (2024)
On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
by: Xiao, Jiancong, et al.
Published: (2024)
by: Xiao, Jiancong, et al.
Published: (2024)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
by: Zhang, Chuheng, et al.
Published: (2024)
by: Zhang, Chuheng, et al.
Published: (2024)
Unifying Stable Optimization and Reference Regularization in RLHF
by: He, Li, et al.
Published: (2026)
by: He, Li, et al.
Published: (2026)
Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Role of Bellman Constraints
by: Xu, Tian, et al.
Published: (2026)
by: Xu, Tian, et al.
Published: (2026)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
by: Gao, Zhaolin, et al.
Published: (2024)
by: Gao, Zhaolin, et al.
Published: (2024)
Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
by: Hou, Zhenyu, et al.
Published: (2024)
by: Hou, Zhenyu, et al.
Published: (2024)
Off-Policy Value-Based Reinforcement Learning for Large Language Models
by: Wang, Peng-Yuan, et al.
Published: (2026)
by: Wang, Peng-Yuan, et al.
Published: (2026)
Distributionally Robust Token Optimization in RLHF
by: Jin, Yeping, et al.
Published: (2026)
by: Jin, Yeping, et al.
Published: (2026)
Provably Efficient Online RLHF with One-Pass Reward Modeling
by: Li, Long-Fei, et al.
Published: (2025)
by: Li, Long-Fei, et al.
Published: (2025)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
WPO: Enhancing RLHF with Weighted Preference Optimization
by: Zhou, Wenxuan, et al.
Published: (2024)
by: Zhou, Wenxuan, et al.
Published: (2024)
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
by: Li, Ziniu, et al.
Published: (2023)
by: Li, Ziniu, et al.
Published: (2023)
Learning to Explore: Policy-Guided Outlier Synthesis for Graph Out-of-Distribution Detection
by: Sun, Li, et al.
Published: (2026)
by: Sun, Li, et al.
Published: (2026)
Factored Causal Representation Learning for Robust Reward Modeling in RLHF
by: Yang, Yupei, et al.
Published: (2026)
by: Yang, Yupei, et al.
Published: (2026)
BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation
by: Jia, Chengxing, et al.
Published: (2024)
by: Jia, Chengxing, et al.
Published: (2024)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
by: Dai, Juntao, et al.
Published: (2025)
by: Dai, Juntao, et al.
Published: (2025)
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
by: Yang, Tianyun, et al.
Published: (2025)
by: Yang, Tianyun, et al.
Published: (2025)
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
by: Mukherjee, Arpan, et al.
Published: (2025)
by: Mukherjee, Arpan, et al.
Published: (2025)
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
by: Huang, Jiawei, et al.
Published: (2025)
by: Huang, Jiawei, et al.
Published: (2025)
Learning a Pessimistic Reward Model in RLHF
by: Xu, Yinglun, et al.
Published: (2025)
by: Xu, Yinglun, et al.
Published: (2025)
Group Robust Preference Optimization in Reward-free RLHF
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
by: Ramesh, Shyam Sundhar, et al.
Published: (2024)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
by: Cen, Shicong, et al.
Published: (2024)
by: Cen, Shicong, et al.
Published: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
by: Hu, Jian, et al.
Published: (2024)
by: Hu, Jian, et al.
Published: (2024)
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
RLHFSpec: Breaking the Efficiency Bottleneck in RLHF Training via Adaptive Drafting
by: Wang, Siqi, et al.
Published: (2025)
by: Wang, Siqi, et al.
Published: (2025)
DPO Meets PPO: Reinforced Token Optimization for RLHF
by: Zhong, Han, et al.
Published: (2024)
by: Zhong, Han, et al.
Published: (2024)
OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
by: Yan, Kaizhuo, et al.
Published: (2025)
by: Yan, Kaizhuo, et al.
Published: (2025)
Active Preference Optimization for Sample Efficient RLHF
by: Das, Nirjhar, et al.
Published: (2024)
by: Das, Nirjhar, et al.
Published: (2024)
Understanding and Alleviating Memory Consumption in RLHF for LLMs
by: Zhou, Jin, et al.
Published: (2024)
by: Zhou, Jin, et al.
Published: (2024)
Mitigating the Alignment Tax of RLHF
by: Lin, Yong, et al.
Published: (2023)
by: Lin, Yong, et al.
Published: (2023)
Tail-Aware Information-Theoretic Generalization for RLHF and SGLD
by: Zhang, Huiming, et al.
Published: (2026)
by: Zhang, Huiming, et al.
Published: (2026)
RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders
by: Yang, Zhongheng, et al.
Published: (2025)
by: Yang, Zhongheng, et al.
Published: (2025)
Similar Items
-
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
by: Xu, Wenyuan, et al.
Published: (2025) -
Dataset Reset Policy Optimization for RLHF
by: Chang, Jonathan D., et al.
Published: (2024) -
Understanding Adversarial Imitation Learning in Small Sample Regime: A Stage-coupled Analysis
by: Xu, Tian, et al.
Published: (2022) -
Efficient Federated RLHF via Zeroth-Order Policy Optimization
by: Wang, Deyi, et al.
Published: (2026) -
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)