Dynamic Planning for LLM-based Graphical User Interface Automation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Shaoqing, Zhang, Zhuosheng, Chen, Kehai, Ma, Xinbei, Yang, Muyun, Zhao, Tiejun, Zhang, Min
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910751799640064
author Zhang, Shaoqing
Zhang, Zhuosheng
Chen, Kehai
Ma, Xinbei
Yang, Muyun
Zhao, Tiejun
Zhang, Min
author_facet Zhang, Shaoqing
Zhang, Zhuosheng
Chen, Kehai
Ma, Xinbei
Yang, Muyun
Zhao, Tiejun
Zhang, Min
contents The advent of large language models (LLMs) has spurred considerable interest in advancing autonomous LLMs-based agents, particularly in intriguing applications within smartphone graphical user interfaces (GUIs). When presented with a task goal, these agents typically emulate human actions within a GUI environment until the task is completed. However, a key challenge lies in devising effective plans to guide action prediction in GUI tasks, though planning have been widely recognized as effective for decomposing complex tasks into a series of steps. Specifically, given the dynamic nature of environmental GUIs following action execution, it is crucial to dynamically adapt plans based on environmental feedback and action history.We show that the widely-used ReAct approach fails due to the excessively long historical dialogues. To address this challenge, we propose a novel approach called Dynamic Planning of Thoughts (D-PoT) for LLM-based GUI agents.D-PoT involves the dynamic adjustment of planning based on the environmental feedback and execution history. Experimental results reveal that the proposed D-PoT significantly surpassed the strong GPT-4V baseline by +12.7% (34.66% $\rightarrow$ 47.36%) in accuracy. The analysis highlights the generality of dynamic planning in different backbone LLMs, as well as the benefits in mitigating hallucinations and adapting to unseen tasks. Code is available at https://github.com/sqzhang-lazy/D-PoT.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00467
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic Planning for LLM-based Graphical User Interface Automation
Zhang, Shaoqing
Zhang, Zhuosheng
Chen, Kehai
Ma, Xinbei
Yang, Muyun
Zhao, Tiejun
Zhang, Min
Artificial Intelligence
Human-Computer Interaction
The advent of large language models (LLMs) has spurred considerable interest in advancing autonomous LLMs-based agents, particularly in intriguing applications within smartphone graphical user interfaces (GUIs). When presented with a task goal, these agents typically emulate human actions within a GUI environment until the task is completed. However, a key challenge lies in devising effective plans to guide action prediction in GUI tasks, though planning have been widely recognized as effective for decomposing complex tasks into a series of steps. Specifically, given the dynamic nature of environmental GUIs following action execution, it is crucial to dynamically adapt plans based on environmental feedback and action history.We show that the widely-used ReAct approach fails due to the excessively long historical dialogues. To address this challenge, we propose a novel approach called Dynamic Planning of Thoughts (D-PoT) for LLM-based GUI agents.D-PoT involves the dynamic adjustment of planning based on the environmental feedback and execution history. Experimental results reveal that the proposed D-PoT significantly surpassed the strong GPT-4V baseline by +12.7% (34.66% $\rightarrow$ 47.36%) in accuracy. The analysis highlights the generality of dynamic planning in different backbone LLMs, as well as the benefits in mitigating hallucinations and adapting to unseen tasks. Code is available at https://github.com/sqzhang-lazy/D-PoT.
title Dynamic Planning for LLM-based Graphical User Interface Automation
topic Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2410.00467