CAAP: Context-Aware Action Planning Prompting to Solve Computer Tasks with Front-End UI Only

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cho, Junhee, Kim, Jihoon, Bae, Daseul, Choo, Jinho, Gwon, Youngjune, Kwon, Yeong-Dae
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915079700611072
author Cho, Junhee
Kim, Jihoon
Bae, Daseul
Choo, Jinho
Gwon, Youngjune
Kwon, Yeong-Dae
author_facet Cho, Junhee
Kim, Jihoon
Bae, Daseul
Choo, Jinho
Gwon, Youngjune
Kwon, Yeong-Dae
contents Software robots have long been used in Robotic Process Automation (RPA) to automate mundane and repetitive computer tasks. With the advent of Large Language Models (LLMs) and their advanced reasoning capabilities, these agents are now able to handle more complex or previously unseen tasks. However, LLM-based automation techniques in recent literature frequently rely on HTML source code for input or application-specific API calls for actions, limiting their applicability to specific environments. We propose an LLM-based agent that mimics human behavior in solving computer tasks. It perceives its environment solely through screenshot images, which are then converted into text for an LLM to process. By leveraging the reasoning capability of the LLM, we eliminate the need for large-scale human demonstration data typically required for model training. The agent only executes keyboard and mouse operations on Graphical User Interface (GUI), removing the need for pre-provided APIs to function. To further enhance the agent's performance in this setting, we propose a novel prompting strategy called Context-Aware Action Planning (CAAP) prompting, which enables the agent to thoroughly examine the task context from multiple perspectives. Our agent achieves an average success rate of 94.5% on MiniWoB++ and an average task score of 62.3 on WebShop, outperforming all previous studies of agents that rely solely on screen images. This method demonstrates potential for broader applications, particularly for tasks requiring coordination across multiple applications on desktops or smartphones, marking a significant advancement in the field of automation agents. Codes and models are accessible at https://github.com/caap-agent/caap-agent.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06947
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CAAP: Context-Aware Action Planning Prompting to Solve Computer Tasks with Front-End UI Only
Cho, Junhee
Kim, Jihoon
Bae, Daseul
Choo, Jinho
Gwon, Youngjune
Kwon, Yeong-Dae
Artificial Intelligence
Human-Computer Interaction
Software robots have long been used in Robotic Process Automation (RPA) to automate mundane and repetitive computer tasks. With the advent of Large Language Models (LLMs) and their advanced reasoning capabilities, these agents are now able to handle more complex or previously unseen tasks. However, LLM-based automation techniques in recent literature frequently rely on HTML source code for input or application-specific API calls for actions, limiting their applicability to specific environments. We propose an LLM-based agent that mimics human behavior in solving computer tasks. It perceives its environment solely through screenshot images, which are then converted into text for an LLM to process. By leveraging the reasoning capability of the LLM, we eliminate the need for large-scale human demonstration data typically required for model training. The agent only executes keyboard and mouse operations on Graphical User Interface (GUI), removing the need for pre-provided APIs to function. To further enhance the agent's performance in this setting, we propose a novel prompting strategy called Context-Aware Action Planning (CAAP) prompting, which enables the agent to thoroughly examine the task context from multiple perspectives. Our agent achieves an average success rate of 94.5% on MiniWoB++ and an average task score of 62.3 on WebShop, outperforming all previous studies of agents that rely solely on screen images. This method demonstrates potential for broader applications, particularly for tasks requiring coordination across multiple applications on desktops or smartphones, marking a significant advancement in the field of automation agents. Codes and models are accessible at https://github.com/caap-agent/caap-agent.
title CAAP: Context-Aware Action Planning Prompting to Solve Computer Tasks with Front-End UI Only
topic Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2406.06947