GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Kangjia, Song, Jiahui, Sha, Leigang, Shen, Haozhan, Chen, Zhi, Zhao, Tiancheng, Liang, Xiubo, Yin, Jianwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915078380453888
author Zhao, Kangjia
Song, Jiahui
Sha, Leigang
Shen, Haozhan
Chen, Zhi
Zhao, Tiancheng
Liang, Xiubo
Yin, Jianwei
author_facet Zhao, Kangjia
Song, Jiahui
Sha, Leigang
Shen, Haozhan
Chen, Zhi
Zhao, Tiancheng
Liang, Xiubo
Yin, Jianwei
contents Nowadays, research on GUI agents is a hot topic in the AI community. However, current research focuses on GUI task automation, limiting the scope of applications in various GUI scenarios. In this paper, we propose a formalized and comprehensive environment to evaluate the entire process of automated GUI Testing (GTArena), offering a fair, standardized environment for consistent operation of diverse multimodal large language models. We divide the testing process into three key subtasks: test intention generation, test task execution, and GUI defect detection, and construct a benchmark dataset based on these to conduct a comprehensive evaluation. It evaluates the performance of different models using three data types: real mobile applications, mobile applications with artificially injected defects, and synthetic data, thoroughly assessing their capabilities in this relevant task. Additionally, we propose a method that helps researchers explore the correlation between the performance of multimodal language large models in specific scenarios and their general capabilities in standard benchmark tests. Experimental results indicate that even the most advanced models struggle to perform well across all sub-tasks of automated GUI Testing, highlighting a significant gap between the current capabilities of Autonomous GUI Testing and its practical, real-world applicability. This gap provides guidance for the future direction of GUI Agent development. Our code is available at https://github.com/ZJU-ACES-ISE/ChatUITest.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18426
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
Zhao, Kangjia
Song, Jiahui
Sha, Leigang
Shen, Haozhan
Chen, Zhi
Zhao, Tiancheng
Liang, Xiubo
Yin, Jianwei
Artificial Intelligence
Nowadays, research on GUI agents is a hot topic in the AI community. However, current research focuses on GUI task automation, limiting the scope of applications in various GUI scenarios. In this paper, we propose a formalized and comprehensive environment to evaluate the entire process of automated GUI Testing (GTArena), offering a fair, standardized environment for consistent operation of diverse multimodal large language models. We divide the testing process into three key subtasks: test intention generation, test task execution, and GUI defect detection, and construct a benchmark dataset based on these to conduct a comprehensive evaluation. It evaluates the performance of different models using three data types: real mobile applications, mobile applications with artificially injected defects, and synthetic data, thoroughly assessing their capabilities in this relevant task. Additionally, we propose a method that helps researchers explore the correlation between the performance of multimodal language large models in specific scenarios and their general capabilities in standard benchmark tests. Experimental results indicate that even the most advanced models struggle to perform well across all sub-tasks of automated GUI Testing, highlighting a significant gap between the current capabilities of Autonomous GUI Testing and its practical, real-world applicability. This gap provides guidance for the future direction of GUI Agent development. Our code is available at https://github.com/ZJU-ACES-ISE/ChatUITest.
title GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
topic Artificial Intelligence
url https://arxiv.org/abs/2412.18426