Efficient Exploration at Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asghari, Seyed Mohammad, Chute, Chris, Dwaracherla, Vikranth, Lu, Xiuyuan, Jafarnia, Mehdi, Minden, Victor, Wen, Zheng, Van Roy, Benjamin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918395306311680
author Asghari, Seyed Mohammad
Chute, Chris
Dwaracherla, Vikranth
Lu, Xiuyuan
Jafarnia, Mehdi
Minden, Victor
Wen, Zheng
Van Roy, Benjamin
author_facet Asghari, Seyed Mohammad
Chute, Chris
Dwaracherla, Vikranth
Lu, Xiuyuan
Jafarnia, Mehdi
Minden, Victor
Wen, Zheng
Van Roy, Benjamin
contents We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates reward and language models as choice data is received. The reward model is fit to the choice data, while the language model is updated by a variation of reinforce, with reinforcement signals provided by the reward model. Several features enable the efficiency gains: a small affirmative nudge added to each reinforcement signal, an epistemic neural network that models reward uncertainty, and information-directed exploration. With Gemma large language models (LLMs), our algorithm matches the performance of offline RLHF trained on 200K labels using fewer than 20K labels, representing more than a 10x gain in data efficiency. Extrapolating from our results, we expect our algorithm trained on 1M labels to match offline RLHF trained on 1B labels. This represents a 1,000x gain. To our knowledge, these are the first results to demonstrate that such large improvements are possible.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17378
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Efficient Exploration at Scale
Asghari, Seyed Mohammad
Chute, Chris
Dwaracherla, Vikranth
Lu, Xiuyuan
Jafarnia, Mehdi
Minden, Victor
Wen, Zheng
Van Roy, Benjamin
Machine Learning
Artificial Intelligence
We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates reward and language models as choice data is received. The reward model is fit to the choice data, while the language model is updated by a variation of reinforce, with reinforcement signals provided by the reward model. Several features enable the efficiency gains: a small affirmative nudge added to each reinforcement signal, an epistemic neural network that models reward uncertainty, and information-directed exploration. With Gemma large language models (LLMs), our algorithm matches the performance of offline RLHF trained on 200K labels using fewer than 20K labels, representing more than a 10x gain in data efficiency. Extrapolating from our results, we expect our algorithm trained on 1M labels to match offline RLHF trained on 1B labels. This represents a 1,000x gain. To our knowledge, these are the first results to demonstrate that such large improvements are possible.
title Efficient Exploration at Scale
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.17378