ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Gao, Difei, Ji, Lei, Bai, Zechen, Ouyang, Mingyu, Li, Peiran, Mao, Dongxing, Wu, Qinchen, Zhang, Weichen, Wang, Peiyi, Guo, Xiangwu, Wang, Hengxu, Zhou, Luowei, Shou, Mike Zheng
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911743991611392
author Gao, Difei
Ji, Lei
Bai, Zechen
Ouyang, Mingyu
Li, Peiran
Mao, Dongxing
Wu, Qinchen
Zhang, Weichen
Wang, Peiyi
Guo, Xiangwu
Wang, Hengxu
Zhou, Luowei
Shou, Mike Zheng
author_facet Gao, Difei
Ji, Lei
Bai, Zechen
Ouyang, Mingyu
Li, Peiran
Mao, Dongxing
Wu, Qinchen
Zhang, Weichen
Wang, Peiyi
Guo, Xiangwu
Wang, Hengxu
Zhou, Luowei
Shou, Mike Zheng
contents Graphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks, thereby boosting human productivity. Existing works leveraging Large Language Model (LLM) or LLM-based AI agents have shown capabilities in automating tasks on Android and Web platforms. However, these tasks are primarily aimed at simple device usage and entertainment operations. This paper presents a novel benchmark, AssistGUI, to evaluate whether models are capable of manipulating the mouse and keyboard on the Windows platform in response to user-requested tasks. We carefully collected a set of 100 tasks from nine widely-used software applications, such as, After Effects and MS Word, each accompanied by the necessary project files for better evaluation. Moreover, we propose an advanced Actor-Critic Embodied Agent framework, which incorporates a sophisticated GUI parser driven by an LLM-agent and an enhanced reasoning mechanism adept at handling lengthy procedural tasks. Our experimental results reveal that our GUI Parser and Reasoning mechanism outshine existing methods in performance. Nevertheless, the potential remains substantial, with the best model attaining only a 46% success rate on our benchmark. We conclude with a thorough analysis of the current methods' limitations, setting the stage for future breakthroughs in this domain.
format Preprint
id arxiv_https___arxiv_org_abs_2312_13108
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
Gao, Difei
Ji, Lei
Bai, Zechen
Ouyang, Mingyu
Li, Peiran
Mao, Dongxing
Wu, Qinchen
Zhang, Weichen
Wang, Peiyi
Guo, Xiangwu
Wang, Hengxu
Zhou, Luowei
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Graphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks, thereby boosting human productivity. Existing works leveraging Large Language Model (LLM) or LLM-based AI agents have shown capabilities in automating tasks on Android and Web platforms. However, these tasks are primarily aimed at simple device usage and entertainment operations. This paper presents a novel benchmark, AssistGUI, to evaluate whether models are capable of manipulating the mouse and keyboard on the Windows platform in response to user-requested tasks. We carefully collected a set of 100 tasks from nine widely-used software applications, such as, After Effects and MS Word, each accompanied by the necessary project files for better evaluation. Moreover, we propose an advanced Actor-Critic Embodied Agent framework, which incorporates a sophisticated GUI parser driven by an LLM-agent and an enhanced reasoning mechanism adept at handling lengthy procedural tasks. Our experimental results reveal that our GUI Parser and Reasoning mechanism outshine existing methods in performance. Nevertheless, the potential remains substantial, with the best model attaining only a 46% success rate on our benchmark. We conclude with a thorough analysis of the current methods' limitations, setting the stage for future breakthroughs in this domain.
title ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.13108