Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Weikai, Jiang, Zhizheng, Liu, Yuxuan, Gao, Pengzhi, Liu, Wei, Luan, Jian, Li, Yuanchun, Liu, Yunxin, Wang, Bin, An, Bo
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917239901388800
author Xu, Weikai
Jiang, Zhizheng
Liu, Yuxuan
Gao, Pengzhi
Liu, Wei
Luan, Jian
Li, Yuanchun
Liu, Yunxin
Wang, Bin
An, Bo
author_facet Xu, Weikai
Jiang, Zhizheng
Liu, Yuxuan
Gao, Pengzhi
Liu, Wei
Luan, Jian
Li, Yuanchun
Liu, Yunxin
Wang, Bin
An, Bo
contents VLM-based mobile agents are increasingly popular due to their capabilities to interact with smartphone GUIs and XML-structured texts and to complete daily tasks. However, existing online benchmarks struggle with obtaining stable reward signals due to dynamic environmental changes. Offline benchmarks evaluate the agents through single-path trajectories, which stands in contrast to the inherently multi-solution characteristics of GUI tasks. Additionally, both types of benchmarks fail to assess whether mobile agents can handle noise or engage in proactive interactions due to a lack of noisy apps or overly full instructions during the evaluation process. To address these limitations, we use a slot-based instruction generation method to construct a more realistic and comprehensive benchmark named Mobile-Bench-v2. Mobile-Bench-v2 includes a common task split, with offline multi-path evaluation to assess the agent's ability to obtain step rewards during task execution. It contains a noisy split based on pop-ups and ads apps, and a contaminated split named AITZ-Noise to formulate a real noisy environment. Furthermore, an ambiguous instruction split with preset Q\&A interactions is released to evaluate the agent's proactive interaction capabilities. We conduct evaluations on these splits using the single-agent framework AppAgent-v1, the multi-agent framework Mobile-Agent-v2, as well as other mobile agents such as UI-Tars and OS-Atlas. Code and data are available at https://huggingface.co/datasets/xwk123/MobileBench-v2.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
Xu, Weikai
Jiang, Zhizheng
Liu, Yuxuan
Gao, Pengzhi
Liu, Wei
Luan, Jian
Li, Yuanchun
Liu, Yunxin
Wang, Bin
An, Bo
Computation and Language
Artificial Intelligence
VLM-based mobile agents are increasingly popular due to their capabilities to interact with smartphone GUIs and XML-structured texts and to complete daily tasks. However, existing online benchmarks struggle with obtaining stable reward signals due to dynamic environmental changes. Offline benchmarks evaluate the agents through single-path trajectories, which stands in contrast to the inherently multi-solution characteristics of GUI tasks. Additionally, both types of benchmarks fail to assess whether mobile agents can handle noise or engage in proactive interactions due to a lack of noisy apps or overly full instructions during the evaluation process. To address these limitations, we use a slot-based instruction generation method to construct a more realistic and comprehensive benchmark named Mobile-Bench-v2. Mobile-Bench-v2 includes a common task split, with offline multi-path evaluation to assess the agent's ability to obtain step rewards during task execution. It contains a noisy split based on pop-ups and ads apps, and a contaminated split named AITZ-Noise to formulate a real noisy environment. Furthermore, an ambiguous instruction split with preset Q\&A interactions is released to evaluate the agent's proactive interaction capabilities. We conduct evaluations on these splits using the single-agent framework AppAgent-v1, the multi-agent framework Mobile-Agent-v2, as well as other mobile agents such as UI-Tars and OS-Atlas. Code and data are available at https://huggingface.co/datasets/xwk123/MobileBench-v2.
title Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.11891