Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Fuxiang, Li, Junyou, Li, Yi-Chen, Zhang, Zongzhang, Yu, Yang, Ye, Deheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911945406283776
author Zhang, Fuxiang
Li, Junyou
Li, Yi-Chen
Zhang, Zongzhang
Yu, Yang
Ye, Deheng
author_facet Zhang, Fuxiang
Li, Junyou
Li, Yi-Chen
Zhang, Zongzhang
Yu, Yang
Ye, Deheng
contents Low sample efficiency is an enduring challenge of reinforcement learning (RL). With the advent of versatile large language models (LLMs), recent works impart common-sense knowledge to accelerate policy learning for RL processes. However, we note that such guidance is often tailored for one specific task but loses generalizability. In this paper, we introduce a framework that harnesses LLMs to extract background knowledge of an environment, which contains general understandings of the entire environment, making various downstream RL tasks benefit from one-time knowledge representation. We ground LLMs by feeding a few pre-collected experiences and requesting them to delineate background knowledge of the environment. Afterward, we represent the output knowledge as potential functions for potential-based reward shaping, which has a good property for maintaining policy optimality from task rewards. We instantiate three variants to prompt LLMs for background knowledge, including writing code, annotating preferences, and assigning goals. Our experiments show that these methods achieve significant sample efficiency improvements in a spectrum of downstream tasks from Minigrid and Crafter domains.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03964
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models
Zhang, Fuxiang
Li, Junyou
Li, Yi-Chen
Zhang, Zongzhang
Yu, Yang
Ye, Deheng
Computation and Language
Machine Learning
Low sample efficiency is an enduring challenge of reinforcement learning (RL). With the advent of versatile large language models (LLMs), recent works impart common-sense knowledge to accelerate policy learning for RL processes. However, we note that such guidance is often tailored for one specific task but loses generalizability. In this paper, we introduce a framework that harnesses LLMs to extract background knowledge of an environment, which contains general understandings of the entire environment, making various downstream RL tasks benefit from one-time knowledge representation. We ground LLMs by feeding a few pre-collected experiences and requesting them to delineate background knowledge of the environment. Afterward, we represent the output knowledge as potential functions for potential-based reward shaping, which has a good property for maintaining policy optimality from task rewards. We instantiate three variants to prompt LLMs for background knowledge, including writing code, annotating preferences, and assigning goals. Our experiments show that these methods achieve significant sample efficiency improvements in a spectrum of downstream tasks from Minigrid and Crafter domains.
title Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.03964