Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Qinqing, Henaff, Mikael, Zhang, Amy, Grover, Aditya, Amos, Brandon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909866401988608
author Zheng, Qinqing
Henaff, Mikael
Zhang, Amy
Grover, Aditya
Amos, Brandon
author_facet Zheng, Qinqing
Henaff, Mikael
Zhang, Amy
Grover, Aditya
Amos, Brandon
contents Automatically synthesizing dense rewards from natural language descriptions is a promising paradigm in reinforcement learning (RL), with applications to sparse reward problems, open-ended exploration, and hierarchical skill design. Recent works have made promising steps by exploiting the prior knowledge of large language models (LLMs). However, these approaches suffer from important limitations: they are either not scalable to problems requiring billions of environment samples, due to requiring LLM annotations for each observation, or they require a diverse offline dataset, which may not exist or be impossible to collect. In this work, we address these limitations through a combination of algorithmic and systems-level contributions. We propose ONI, a distributed architecture that simultaneously learns an RL policy and an intrinsic reward function using LLM feedback. Our approach annotates the agent's collected experience via an asynchronous LLM server, which is then distilled into an intrinsic reward model. We explore a range of algorithmic choices for reward modeling with varying complexity, including hashing, classification, and ranking models. Our approach achieves state-of-the-art performance across a range of challenging tasks from the NetHack Learning Environment, while removing the need for large offline datasets required by prior work. We make our code available at https://github.com/facebookresearch/oni.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23022
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
Zheng, Qinqing
Henaff, Mikael
Zhang, Amy
Grover, Aditya
Amos, Brandon
Machine Learning
Artificial Intelligence
Computation and Language
Robotics
Automatically synthesizing dense rewards from natural language descriptions is a promising paradigm in reinforcement learning (RL), with applications to sparse reward problems, open-ended exploration, and hierarchical skill design. Recent works have made promising steps by exploiting the prior knowledge of large language models (LLMs). However, these approaches suffer from important limitations: they are either not scalable to problems requiring billions of environment samples, due to requiring LLM annotations for each observation, or they require a diverse offline dataset, which may not exist or be impossible to collect. In this work, we address these limitations through a combination of algorithmic and systems-level contributions. We propose ONI, a distributed architecture that simultaneously learns an RL policy and an intrinsic reward function using LLM feedback. Our approach annotates the agent's collected experience via an asynchronous LLM server, which is then distilled into an intrinsic reward model. We explore a range of algorithmic choices for reward modeling with varying complexity, including hashing, classification, and ranking models. Our approach achieves state-of-the-art performance across a range of challenging tasks from the NetHack Learning Environment, while removing the need for large offline datasets required by prior work. We make our code available at https://github.com/facebookresearch/oni.
title Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
topic Machine Learning
Artificial Intelligence
Computation and Language
Robotics
url https://arxiv.org/abs/2410.23022