Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fei, Zhaoye, Ji, Li, Wang, Siyin, Shi, Junhao, Gong, Jingjing, Qiu, Xipeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913917385572352
author Fei, Zhaoye
Ji, Li
Wang, Siyin
Shi, Junhao
Gong, Jingjing
Qiu, Xipeng
author_facet Fei, Zhaoye
Ji, Li
Wang, Siyin
Shi, Junhao
Gong, Jingjing
Qiu, Xipeng
contents Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they face significant challenges in embodied task planning scenarios that require continuous environmental understanding and action generation. Existing approaches generate open-loop action scripts based on static knowledge, making it difficult to learn causal relationships between actions and environmental feedback, particularly in partially observable environments. We introduce Embodied Planner-R1, a novel outcome-driven reinforcement learning framework that enables LLMs to develop interactive capabilities through autonomous exploration with minimal supervision. Our framework incorporates three key innovations: (1) Without human annotations, we employ pure reinforcement learning with group rollout, incorporating in-environment interaction through parallel exploration; (2) completion-driven sparse reward; and (3) Interactive Policy Optimization (IPO) for efficient learning from grouped trajectories. Across two challenging text-based Embodied planning benchmarks, Embodied Planner-R1 achieves impressive completion rates of 97.78% on ALFWorld and 79.92% on ScienceWorld, surpassing prior methods by a large margin, and suffers only a -3.66% drop in previously unseen environments, evidencing strong generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
Fei, Zhaoye
Ji, Li
Wang, Siyin
Shi, Junhao
Gong, Jingjing
Qiu, Xipeng
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they face significant challenges in embodied task planning scenarios that require continuous environmental understanding and action generation. Existing approaches generate open-loop action scripts based on static knowledge, making it difficult to learn causal relationships between actions and environmental feedback, particularly in partially observable environments. We introduce Embodied Planner-R1, a novel outcome-driven reinforcement learning framework that enables LLMs to develop interactive capabilities through autonomous exploration with minimal supervision. Our framework incorporates three key innovations: (1) Without human annotations, we employ pure reinforcement learning with group rollout, incorporating in-environment interaction through parallel exploration; (2) completion-driven sparse reward; and (3) Interactive Policy Optimization (IPO) for efficient learning from grouped trajectories. Across two challenging text-based Embodied planning benchmarks, Embodied Planner-R1 achieves impressive completion rates of 97.78% on ALFWorld and 79.92% on ScienceWorld, surpassing prior methods by a large margin, and suffers only a -3.66% drop in previously unseen environments, evidencing strong generalization.
title Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.23127