Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911945406283776 |
|---|---|
| author | Zhang, Fuxiang Li, Junyou Li, Yi-Chen Zhang, Zongzhang Yu, Yang Ye, Deheng |
| author_facet | Zhang, Fuxiang Li, Junyou Li, Yi-Chen Zhang, Zongzhang Yu, Yang Ye, Deheng |
| contents | Low sample efficiency is an enduring challenge of reinforcement learning (RL). With the advent of versatile large language models (LLMs), recent works impart common-sense knowledge to accelerate policy learning for RL processes. However, we note that such guidance is often tailored for one specific task but loses generalizability. In this paper, we introduce a framework that harnesses LLMs to extract background knowledge of an environment, which contains general understandings of the entire environment, making various downstream RL tasks benefit from one-time knowledge representation. We ground LLMs by feeding a few pre-collected experiences and requesting them to delineate background knowledge of the environment. Afterward, we represent the output knowledge as potential functions for potential-based reward shaping, which has a good property for maintaining policy optimality from task rewards. We instantiate three variants to prompt LLMs for background knowledge, including writing code, annotating preferences, and assigning goals. Our experiments show that these methods achieve significant sample efficiency improvements in a spectrum of downstream tasks from Minigrid and Crafter domains. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_03964 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models Zhang, Fuxiang Li, Junyou Li, Yi-Chen Zhang, Zongzhang Yu, Yang Ye, Deheng Computation and Language Machine Learning Low sample efficiency is an enduring challenge of reinforcement learning (RL). With the advent of versatile large language models (LLMs), recent works impart common-sense knowledge to accelerate policy learning for RL processes. However, we note that such guidance is often tailored for one specific task but loses generalizability. In this paper, we introduce a framework that harnesses LLMs to extract background knowledge of an environment, which contains general understandings of the entire environment, making various downstream RL tasks benefit from one-time knowledge representation. We ground LLMs by feeding a few pre-collected experiences and requesting them to delineate background knowledge of the environment. Afterward, we represent the output knowledge as potential functions for potential-based reward shaping, which has a good property for maintaining policy optimality from task rewards. We instantiate three variants to prompt LLMs for background knowledge, including writing code, annotating preferences, and assigning goals. Our experiments show that these methods achieve significant sample efficiency improvements in a spectrum of downstream tasks from Minigrid and Crafter domains. |
| title | Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2407.03964 |