Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xie, Guanwen, Xu, Jingzehua, Yang, Yiyuan, Ding, Yimian, Zhang, Shuai
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913133077987328
author Xie, Guanwen
Xu, Jingzehua
Yang, Yiyuan
Ding, Yimian
Zhang, Shuai
author_facet Xie, Guanwen
Xu, Jingzehua
Yang, Yiyuan
Ding, Yimian
Zhang, Shuai
contents Achieving the effective design and improvement of reward functions in reinforcement learning (RL) tasks with complex custom environments and multiple requirements presents considerable challenges. In this paper, we propose ERFSL, an efficient reward function searcher using LLMs, which enables LLMs to be effective white-box searchers and highlights their advanced semantic understanding capabilities. Specifically, we generate reward components for each numerically explicit user requirement and employ a reward critic to identify the correct code form. Then, LLMs assign weights to the reward components to balance their values and iteratively adjust the weights without ambiguity and redundant adjustments by flexibly adopting directional mutation and crossover strategies, similar to genetic algorithms, based on the context provided by the training log analyzer. We applied the framework to a customized data collection RL task without direct human feedback or reward examples (zero-shot learning). The reward critic successfully corrects the reward code with only one feedback instance for each requirement, effectively preventing unrectifiable errors. The initialization of weights enables the acquisition of different reward functions within the Pareto solution set without the need for weight search. Even in cases where a weight is 500 times off, on average, only 5.2 iterations are needed to meet user requirements. The ERFSL also works well with most prompts utilizing GPT-4o mini, as we decompose the weight searching process to reduce the requirement for numerical and long-context understanding capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2409_02428
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement
Xie, Guanwen
Xu, Jingzehua
Yang, Yiyuan
Ding, Yimian
Zhang, Shuai
Machine Learning
Artificial Intelligence
Computation and Language
Systems and Control
Achieving the effective design and improvement of reward functions in reinforcement learning (RL) tasks with complex custom environments and multiple requirements presents considerable challenges. In this paper, we propose ERFSL, an efficient reward function searcher using LLMs, which enables LLMs to be effective white-box searchers and highlights their advanced semantic understanding capabilities. Specifically, we generate reward components for each numerically explicit user requirement and employ a reward critic to identify the correct code form. Then, LLMs assign weights to the reward components to balance their values and iteratively adjust the weights without ambiguity and redundant adjustments by flexibly adopting directional mutation and crossover strategies, similar to genetic algorithms, based on the context provided by the training log analyzer. We applied the framework to a customized data collection RL task without direct human feedback or reward examples (zero-shot learning). The reward critic successfully corrects the reward code with only one feedback instance for each requirement, effectively preventing unrectifiable errors. The initialization of weights enables the acquisition of different reward functions within the Pareto solution set without the need for weight search. Even in cases where a weight is 500 times off, on average, only 5.2 iterations are needed to meet user requirements. The ERFSL also works well with most prompts utilizing GPT-4o mini, as we decompose the weight searching process to reduce the requirement for numerical and long-context understanding capabilities.
title Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement
topic Machine Learning
Artificial Intelligence
Computation and Language
Systems and Control
url https://arxiv.org/abs/2409.02428