Saved in:
Bibliographic Details
Main Author: Wu, Xiefeng
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2405.03341
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929356828311552
author Wu, Xiefeng
author_facet Wu, Xiefeng
contents Q-learning excels in learning from feedback within sequential decision-making tasks but often requires extensive sampling to achieve significant improvements. While reward shaping can enhance learning efficiency, non-potential-based methods introduce biases that affect performance, and potential-based reward shaping, though unbiased, lacks the ability to provide heuristics for state-action pairs, limiting its effectiveness in complex environments. Large language models (LLMs) can achieve zero-shot learning for simpler tasks, but they suffer from low inference speeds and occasional hallucinations. To address these challenges, we propose \textbf{LLM-guided Q-learning}, a framework that leverages LLMs as heuristics to aid in learning the Q-function for reinforcement learning. Our theoretical analysis demonstrates that this approach adapts to hallucinations, improves sample efficiency, and avoids biasing final performance. Experimental results show that our algorithm is general, robust, and capable of preventing ineffective exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03341
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Q-Learning with Large Language Model Heuristics
Wu, Xiefeng
Machine Learning
Artificial Intelligence
Q-learning excels in learning from feedback within sequential decision-making tasks but often requires extensive sampling to achieve significant improvements. While reward shaping can enhance learning efficiency, non-potential-based methods introduce biases that affect performance, and potential-based reward shaping, though unbiased, lacks the ability to provide heuristics for state-action pairs, limiting its effectiveness in complex environments. Large language models (LLMs) can achieve zero-shot learning for simpler tasks, but they suffer from low inference speeds and occasional hallucinations. To address these challenges, we propose \textbf{LLM-guided Q-learning}, a framework that leverages LLMs as heuristics to aid in learning the Q-function for reinforcement learning. Our theoretical analysis demonstrates that this approach adapts to hallucinations, improves sample efficiency, and avoids biasing final performance. Experimental results show that our algorithm is general, robust, and capable of preventing ineffective exploration.
title Enhancing Q-Learning with Large Language Model Heuristics
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.03341