Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Qining, Ying, Lei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reinforcement Learning from Human Feedback without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis
por: Zhang, Qining, et al.
Publicado: (2024)
por: Zhang, Qining, et al.
Publicado: (2024)
Efficient Federated RLHF via Zeroth-Order Policy Optimization
por: Wang, Deyi, et al.
Publicado: (2026)
por: Wang, Deyi, et al.
Publicado: (2026)
Provable Reinforcement Learning from Human Feedback with an Unknown Link Function
por: Zhang, Qining, et al.
Publicado: (2025)
por: Zhang, Qining, et al.
Publicado: (2025)
Fast and Regret Optimal Best Arm Identification: Fundamental Limits and Low-Complexity Algorithms
por: Zhang, Qining, et al.
Publicado: (2023)
por: Zhang, Qining, et al.
Publicado: (2023)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback
por: Liu, Qiang, et al.
Publicado: (2026)
por: Liu, Qiang, et al.
Publicado: (2026)
Zeroth-Order Optimization Meets Human Feedback: Provable Learning via Ranking Oracles
por: Tang, Zhiwei, et al.
Publicado: (2023)
por: Tang, Zhiwei, et al.
Publicado: (2023)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
por: Ackermann, Johannes, et al.
Publicado: (2025)
por: Ackermann, Johannes, et al.
Publicado: (2025)
Dense Reward for Free in Reinforcement Learning from Human Feedback
por: Chan, Alex J., et al.
Publicado: (2024)
por: Chan, Alex J., et al.
Publicado: (2024)
Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
por: Zhai, Yuanzhao, et al.
Publicado: (2023)
por: Zhai, Yuanzhao, et al.
Publicado: (2023)
Cost Aware Best Arm Identification
por: Kanarios, Kellen, et al.
Publicado: (2024)
por: Kanarios, Kellen, et al.
Publicado: (2024)
Data-Free Black-Box Federated Learning via Zeroth-Order Gradient Estimation
por: Ma, Xinge, et al.
Publicado: (2025)
por: Ma, Xinge, et al.
Publicado: (2025)
Ancestral Reinforcement Learning: Unifying Zeroth-Order Optimization and Genetic Algorithms for Reinforcement Learning
por: Nakashima, So, et al.
Publicado: (2024)
por: Nakashima, So, et al.
Publicado: (2024)
Powering Up Zeroth-Order Training via Subspace Gradient Orthogonalization
por: Lang, Yicheng, et al.
Publicado: (2026)
por: Lang, Yicheng, et al.
Publicado: (2026)
Policy Learning from Large Vision-Language Model Feedback without Reward Modeling
por: Luu, Tung M., et al.
Publicado: (2025)
por: Luu, Tung M., et al.
Publicado: (2025)
LLM Zeroth-Order Fine-Tuning is an Inference Workload
por: Li, Zelin, et al.
Publicado: (2026)
por: Li, Zelin, et al.
Publicado: (2026)
Learning Dynamics of Zeroth-Order Optimization: A Kernel Perspective
por: Li, Zhe, et al.
Publicado: (2026)
por: Li, Zhe, et al.
Publicado: (2026)
Zeroth-Order Hard-Thresholding: Gradient Error vs. Expansivity
por: de Vazelhes, William, et al.
Publicado: (2022)
por: de Vazelhes, William, et al.
Publicado: (2022)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
por: Gupta, Devansh, et al.
Publicado: (2025)
por: Gupta, Devansh, et al.
Publicado: (2025)
Privacy-Preserving Reinforcement Learning from Human Feedback via Decoupled Reward Modeling
por: Cho, Young Hyun, et al.
Publicado: (2026)
por: Cho, Young Hyun, et al.
Publicado: (2026)
Zeroth-Order Optimization is Secretly Single-Step Policy Optimization
por: Qiu, Junbin, et al.
Publicado: (2025)
por: Qiu, Junbin, et al.
Publicado: (2025)
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
por: Zhang, Shun, et al.
Publicado: (2024)
por: Zhang, Shun, et al.
Publicado: (2024)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
por: Chakraborty, Souradip, et al.
Publicado: (2023)
por: Chakraborty, Souradip, et al.
Publicado: (2023)
Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs
por: Zhang, Kairun, et al.
Publicado: (2025)
por: Zhang, Kairun, et al.
Publicado: (2025)
Second-Order Fine-Tuning without Pain for LLMs:A Hessian Informed Zeroth-Order Optimizer
por: Zhao, Yanjun, et al.
Publicado: (2024)
por: Zhao, Yanjun, et al.
Publicado: (2024)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
por: Xiao, Teng, et al.
Publicado: (2024)
por: Xiao, Teng, et al.
Publicado: (2024)
Gradient Compressed Sensing: A Query-Efficient Gradient Estimator for High-Dimensional Zeroth-Order Optimization
por: Qiu, Ruizhong, et al.
Publicado: (2024)
por: Qiu, Ruizhong, et al.
Publicado: (2024)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
por: Xiao, Minheng, et al.
Publicado: (2024)
por: Xiao, Minheng, et al.
Publicado: (2024)
ZOTTA: Test-Time Adaptation with Gradient-Free Zeroth-Order Optimization
por: Zhang, Ronghao, et al.
Publicado: (2026)
por: Zhang, Ronghao, et al.
Publicado: (2026)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
por: Ma, Shaocong, et al.
Publicado: (2025)
por: Ma, Shaocong, et al.
Publicado: (2025)
Reinforcement Learning from Human Feedback
por: Lambert, Nathan
Publicado: (2025)
por: Lambert, Nathan
Publicado: (2025)
Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered
por: Liu, Sijia, et al.
Publicado: (2026)
por: Liu, Sijia, et al.
Publicado: (2026)
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
por: Wan, Zhenglin, et al.
Publicado: (2025)
por: Wan, Zhenglin, et al.
Publicado: (2025)
Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
por: Chaudhari, Shreyas, et al.
Publicado: (2025)
por: Chaudhari, Shreyas, et al.
Publicado: (2025)
Obtaining Lower Query Complexities through Lightweight Zeroth-Order Proximal Gradient Algorithms
por: Gu, Bin, et al.
Publicado: (2024)
por: Gu, Bin, et al.
Publicado: (2024)
Distributed Zeroth-Order Policy Gradient for Networked Multi-agent Reinforcement Learning from Human Feedback
por: Dai, Pengcheng, et al.
Publicado: (2026)
por: Dai, Pengcheng, et al.
Publicado: (2026)
Robust Reinforcement Learning from Corrupted Human Feedback
por: Bukharin, Alexander, et al.
Publicado: (2024)
por: Bukharin, Alexander, et al.
Publicado: (2024)
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
por: Chakraborty, Souradip, et al.
Publicado: (2023)
por: Chakraborty, Souradip, et al.
Publicado: (2023)
TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback
por: Pang, Lei, et al.
Publicado: (2025)
por: Pang, Lei, et al.
Publicado: (2025)
Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
Ejemplares similares
-
Reinforcement Learning from Human Feedback without Reward Inference: Model-Free Algorithm and Instance-Dependent Analysis
por: Zhang, Qining, et al.
Publicado: (2024) -
Efficient Federated RLHF via Zeroth-Order Policy Optimization
por: Wang, Deyi, et al.
Publicado: (2026) -
Provable Reinforcement Learning from Human Feedback with an Unknown Link Function
por: Zhang, Qining, et al.
Publicado: (2025) -
Fast and Regret Optimal Best Arm Identification: Fundamental Limits and Low-Complexity Algorithms
por: Zhang, Qining, et al.
Publicado: (2023) -
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)