Greedy Sampling Is Provably Efficient for RLHF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Di, Shi, Chengshuai, Yang, Jing, Shen, Cong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909874045059072
author Wu, Di
Shi, Chengshuai
Yang, Jing
Shen, Cong
author_facet Wu, Di
Shi, Chengshuai
Yang, Jing
Shen, Cong
contents Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique for post-training large language models. Despite its empirical success, the theoretical understanding of RLHF is still limited, as learning the KL-regularized target with only preference feedback poses additional challenges compared with canonical RL. Existing works mostly study the reward-based Bradley-Terry (BT) preference model, and extend classical designs utilizing optimism or pessimism. This work, instead, considers the general preference model (whose practical relevance has been observed recently) and obtains performance guarantees with major, order-wise improvements over existing ones. Surprisingly, these results are derived from algorithms that directly use the empirical estimates (i.e., greedy sampling), as opposed to constructing optimistic or pessimistic estimates in previous works. This insight has a deep root in the unique structural property of the optimal policy class under the KL-regularized target, and we further specialize it to the BT model, highlighting the surprising sufficiency of greedy sampling in RLHF.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24700
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Greedy Sampling Is Provably Efficient for RLHF
Wu, Di
Shi, Chengshuai
Yang, Jing
Shen, Cong
Machine Learning
Artificial Intelligence
Information Theory
Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique for post-training large language models. Despite its empirical success, the theoretical understanding of RLHF is still limited, as learning the KL-regularized target with only preference feedback poses additional challenges compared with canonical RL. Existing works mostly study the reward-based Bradley-Terry (BT) preference model, and extend classical designs utilizing optimism or pessimism. This work, instead, considers the general preference model (whose practical relevance has been observed recently) and obtains performance guarantees with major, order-wise improvements over existing ones. Surprisingly, these results are derived from algorithms that directly use the empirical estimates (i.e., greedy sampling), as opposed to constructing optimistic or pessimistic estimates in previous works. This insight has a deep root in the unique structural property of the optimal policy class under the KL-regularized target, and we further specialize it to the BT model, highlighting the surprising sufficiency of greedy sampling in RLHF.
title Greedy Sampling Is Provably Efficient for RLHF
topic Machine Learning
Artificial Intelligence
Information Theory
url https://arxiv.org/abs/2510.24700