Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zambrano, Alejandra, Marjanovic, Sara Vera, Kerboua, Imene, Lù, Xing Han, Kosseim, Leila
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914614237724672
author Zambrano, Alejandra
Marjanovic, Sara Vera
Kerboua, Imene
Lù, Xing Han
Kosseim, Leila
author_facet Zambrano, Alejandra
Marjanovic, Sara Vera
Kerboua, Imene
Lù, Xing Han
Kosseim, Leila
contents Despite recent advances, LLM-based web agents still struggle with limited exploration, omission of critical steps, and sensitivity to task constraints. Prior work suggests that many of these failures stem from weaknesses in planning, yet the impact of alternative natural language plan representation remains unexplored. To address this, we introduce PlanAhead, a static planner-executor framework that evaluates the impact of plan representation in agent performance. We first automatically categorize WebArena tasks into 3 difficulty levels, enabling consistent difficulty grading without human annotation. Then we systematically evaluate 4 different plan representations on the tasks categorized as hard: sequential subgoals, narrative, pseudocode, and checklist; across different families of multimodal LLM powered agents (OpenAI, Alibaba, and Google). To account for stochastic variability, we introduce two novel evaluation metrics: Achievement Rate (AR) and Solved-Task Consistency (STC). Our results show that both, the plan formulation and the underlying LLM generating the plan, significantly influence web-agent robustness and task success.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29927
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents
Zambrano, Alejandra
Marjanovic, Sara Vera
Kerboua, Imene
Lù, Xing Han
Kosseim, Leila
Computation and Language
Artificial Intelligence
Machine Learning
Despite recent advances, LLM-based web agents still struggle with limited exploration, omission of critical steps, and sensitivity to task constraints. Prior work suggests that many of these failures stem from weaknesses in planning, yet the impact of alternative natural language plan representation remains unexplored. To address this, we introduce PlanAhead, a static planner-executor framework that evaluates the impact of plan representation in agent performance. We first automatically categorize WebArena tasks into 3 difficulty levels, enabling consistent difficulty grading without human annotation. Then we systematically evaluate 4 different plan representations on the tasks categorized as hard: sequential subgoals, narrative, pseudocode, and checklist; across different families of multimodal LLM powered agents (OpenAI, Alibaba, and Google). To account for stochastic variability, we introduce two novel evaluation metrics: Achievement Rate (AR) and Solved-Task Consistency (STC). Our results show that both, the plan formulation and the underlying LLM generating the plan, significantly influence web-agent robustness and task success.
title Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.29927