AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Jiahui, Hua, Zhichao, Xia, Yubin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912601988923392
author Sun, Jiahui
Hua, Zhichao
Xia, Yubin
author_facet Sun, Jiahui
Hua, Zhichao
Xia, Yubin
contents Comprehensive evaluation of mobile agents can significantly advance their development and real-world applicability. However, existing benchmarks lack practicality and scalability due to the extensive manual effort in defining task reward signals and implementing evaluation codes. We propose AutoEval, an evaluation framework which tests mobile agents without any manual effort. Our approach designs a UI state change representation which can be used to automatically generate task reward signals, and employs a Judge System for autonomous evaluation. Evaluation shows AutoEval can automatically generate reward signals with high correlation to human-annotated signals, and achieve high accuracy (up to 94%) in autonomous evaluation comparable to human evaluation. Finally, we evaluate state-of-the-art mobile agents using our framework, providing insights into their performance and limitations.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02403
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents
Sun, Jiahui
Hua, Zhichao
Xia, Yubin
Artificial Intelligence
Comprehensive evaluation of mobile agents can significantly advance their development and real-world applicability. However, existing benchmarks lack practicality and scalability due to the extensive manual effort in defining task reward signals and implementing evaluation codes. We propose AutoEval, an evaluation framework which tests mobile agents without any manual effort. Our approach designs a UI state change representation which can be used to automatically generate task reward signals, and employs a Judge System for autonomous evaluation. Evaluation shows AutoEval can automatically generate reward signals with high correlation to human-annotated signals, and achieve high accuracy (up to 94%) in autonomous evaluation comparable to human evaluation. Finally, we evaluate state-of-the-art mobile agents using our framework, providing insights into their performance and limitations.
title AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2503.02403