Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Erdogan, Lutfi Eren, Lee, Nicholas, Kim, Sehoon, Moon, Suhong, Furuta, Hiroki, Anumanchipalli, Gopala, Keutzer, Kurt, Gholami, Amir
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911087712010240
author Erdogan, Lutfi Eren
Lee, Nicholas
Kim, Sehoon
Moon, Suhong
Furuta, Hiroki
Anumanchipalli, Gopala
Keutzer, Kurt
Gholami, Amir
author_facet Erdogan, Lutfi Eren
Lee, Nicholas
Kim, Sehoon
Moon, Suhong
Furuta, Hiroki
Anumanchipalli, Gopala
Keutzer, Kurt
Gholami, Amir
contents Large language models (LLMs) have shown remarkable advancements in enabling language agents to tackle simple tasks. However, applying them for complex, multi-step, long-horizon tasks remains a challenge. Recent work have found success by separating high-level planning from low-level execution, which enables the model to effectively balance high-level planning objectives and low-level execution details. However, generating accurate plans remains difficult since LLMs are not inherently trained for this task. To address this, we propose Plan-and-Act, a novel framework that incorporates explicit planning into LLM-based agents and introduces a scalable method to enhance plan generation through a novel synthetic data generation method. Plan-and-Act consists of a Planner model which generates structured, high-level plans to achieve user goals, and an Executor model that translates these plans into environment-specific actions. To train the Planner effectively, we introduce a synthetic data generation method that annotates ground-truth trajectories with feasible plans, augmented with diverse and extensive examples to enhance generalization. We evaluate Plan-and-Act using web navigation as a representative long-horizon planning environment, demonstrating a state-of-the-art 57.58% success rate on the WebArena-Lite benchmark as well as a text-only state-of-the-art 81.36% success rate on WebVoyager.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
Erdogan, Lutfi Eren
Lee, Nicholas
Kim, Sehoon
Moon, Suhong
Furuta, Hiroki
Anumanchipalli, Gopala
Keutzer, Kurt
Gholami, Amir
Computation and Language
Large language models (LLMs) have shown remarkable advancements in enabling language agents to tackle simple tasks. However, applying them for complex, multi-step, long-horizon tasks remains a challenge. Recent work have found success by separating high-level planning from low-level execution, which enables the model to effectively balance high-level planning objectives and low-level execution details. However, generating accurate plans remains difficult since LLMs are not inherently trained for this task. To address this, we propose Plan-and-Act, a novel framework that incorporates explicit planning into LLM-based agents and introduces a scalable method to enhance plan generation through a novel synthetic data generation method. Plan-and-Act consists of a Planner model which generates structured, high-level plans to achieve user goals, and an Executor model that translates these plans into environment-specific actions. To train the Planner effectively, we introduce a synthetic data generation method that annotates ground-truth trajectories with feasible plans, augmented with diverse and extensive examples to enhance generalization. We evaluate Plan-and-Act using web navigation as a representative long-horizon planning environment, demonstrating a state-of-the-art 57.58% success rate on the WebArena-Lite benchmark as well as a text-only state-of-the-art 81.36% success rate on WebVoyager.
title Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
topic Computation and Language
url https://arxiv.org/abs/2503.09572