Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Zhicheng, Wang, Ziyan, Du, Yali, Fang, Fei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2506.20061
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913911713824768
author Zhang, Zhicheng
Wang, Ziyan
Du, Yali
Fang, Fei
author_facet Zhang, Zhicheng
Wang, Ziyan
Du, Yali
Fang, Fei
contents Developing effective instruction-following policies in reinforcement learning remains challenging due to the reliance on extensive human-labeled instruction datasets and the difficulty of learning from sparse rewards. In this paper, we propose a novel approach that leverages the capabilities of large language models (LLMs) to automatically generate open-ended instructions retrospectively from previously collected agent trajectories. Our core idea is to employ LLMs to relabel unsuccessful trajectories by identifying meaningful subtasks the agent has implicitly accomplished, thereby enriching the agent's training data and substantially alleviating reliance on human annotations. Through this open-ended instruction relabeling, we efficiently learn a unified instruction-following policy capable of handling diverse tasks within a single policy. We empirically evaluate our proposed method in the challenging Craftax environment, demonstrating clear improvements in sample efficiency, instruction coverage, and overall policy performance compared to state-of-the-art baselines. Our results highlight the effectiveness of utilizing LLM-guided open-ended instruction relabeling to enhance instruction-following reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20061
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
Zhang, Zhicheng
Wang, Ziyan
Du, Yali
Fang, Fei
Machine Learning
Computation and Language
Developing effective instruction-following policies in reinforcement learning remains challenging due to the reliance on extensive human-labeled instruction datasets and the difficulty of learning from sparse rewards. In this paper, we propose a novel approach that leverages the capabilities of large language models (LLMs) to automatically generate open-ended instructions retrospectively from previously collected agent trajectories. Our core idea is to employ LLMs to relabel unsuccessful trajectories by identifying meaningful subtasks the agent has implicitly accomplished, thereby enriching the agent's training data and substantially alleviating reliance on human annotations. Through this open-ended instruction relabeling, we efficiently learn a unified instruction-following policy capable of handling diverse tasks within a single policy. We empirically evaluate our proposed method in the challenging Craftax environment, demonstrating clear improvements in sample efficiency, instruction coverage, and overall policy performance compared to state-of-the-art baselines. Our results highlight the effectiveness of utilizing LLM-guided open-ended instruction relabeling to enhance instruction-following reinforcement learning.
title Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2506.20061