WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jingbo, Hou, Bairu, Wei, Wei, Chang, Shiyu, Bao, Yujia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918156393512960
author Yang, Jingbo
Hou, Bairu
Wei, Wei
Chang, Shiyu
Bao, Yujia
author_facet Yang, Jingbo
Hou, Bairu
Wei, Wei
Chang, Shiyu
Bao, Yujia
contents Large language model (LLM) agents are becoming competent at straightforward web tasks, such as opening an item page or submitting a form, but still struggle with objectives that require long horizon navigation, large scale information extraction, and reasoning under constraints. We present WebDART, a general framework that enables a single LLM to handle such complex chores. WebDART (i) dynamically decomposes each objective into three focused subtasks: navigation, information extraction, and execution, so the model concentrates on one skill at a time, and (ii) continuously replans the decomposition as new webpages are revealed, taking advantage of newly discovered filters or shortcuts and avoiding redundant exploration. Evaluated on WebChoreArena, WebDART lifts success rates by up to 13.7 percentage points over previous SOTA agents, while matching their performance on the easier WebArena suite and completing tasks with up to 14.7 fewer navigation steps.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06587
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
Yang, Jingbo
Hou, Bairu
Wei, Wei
Chang, Shiyu
Bao, Yujia
Artificial Intelligence
Large language model (LLM) agents are becoming competent at straightforward web tasks, such as opening an item page or submitting a form, but still struggle with objectives that require long horizon navigation, large scale information extraction, and reasoning under constraints. We present WebDART, a general framework that enables a single LLM to handle such complex chores. WebDART (i) dynamically decomposes each objective into three focused subtasks: navigation, information extraction, and execution, so the model concentrates on one skill at a time, and (ii) continuously replans the decomposition as new webpages are revealed, taking advantage of newly discovered filters or shortcuts and avoiding redundant exploration. Evaluated on WebChoreArena, WebDART lifts success rates by up to 13.7 percentage points over previous SOTA agents, while matching their performance on the easier WebArena suite and completing tasks with up to 14.7 fewer navigation steps.
title WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
topic Artificial Intelligence
url https://arxiv.org/abs/2510.06587