Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Duong, Thang, Yang, Minglai, Zhang, Chicheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908367029534720
author Duong, Thang
Yang, Minglai
Zhang, Chicheng
author_facet Duong, Thang
Yang, Minglai
Zhang, Chicheng
contents We investigate the usage of Large Language Model (LLM) in collecting high-quality data to warm-start Reinforcement Learning (RL) algorithms for learning in some classical Markov Decision Process (MDP) environments. In this work, we focus on using LLM to generate an off-policy dataset that sufficiently covers state-actions visited by optimal policies, then later using an RL algorithm to explore the environment and improve the policy suggested by the LLM. Our algorithm, LORO, can both converge to an optimal policy and have a high sample efficiency thanks to the LLM's good starting policy. On multiple OpenAI Gym environments, such as CartPole and Pendulum, we empirically demonstrate that LORO outperforms baseline algorithms such as pure LLM-based policies, pure RL, and a naive combination of the two, achieving up to $4 \times$ the cumulative rewards of the pure RL baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM
Duong, Thang
Yang, Minglai
Zhang, Chicheng
Machine Learning
We investigate the usage of Large Language Model (LLM) in collecting high-quality data to warm-start Reinforcement Learning (RL) algorithms for learning in some classical Markov Decision Process (MDP) environments. In this work, we focus on using LLM to generate an off-policy dataset that sufficiently covers state-actions visited by optimal policies, then later using an RL algorithm to explore the environment and improve the policy suggested by the LLM. Our algorithm, LORO, can both converge to an optimal policy and have a high sample efficiency thanks to the LLM's good starting policy. On multiple OpenAI Gym environments, such as CartPole and Pendulum, we empirically demonstrate that LORO outperforms baseline algorithms such as pure LLM-based policies, pure RL, and a naive combination of the two, achieving up to $4 \times$ the cumulative rewards of the pure RL baseline.
title Improving the Data-efficiency of Reinforcement Learning by Warm-starting with LLM
topic Machine Learning
url https://arxiv.org/abs/2505.10861