MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Im, Youngmin, Jo, Byeongung, Wi, Jaeyoung, Baek, Seungwoo, Min, Tae Hoon, Lee, Joo Hyung, Oh, Sangeun, Shin, Insik, Lee, Sunjae
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910278231261184
author Im, Youngmin
Jo, Byeongung
Wi, Jaeyoung
Baek, Seungwoo
Min, Tae Hoon
Lee, Joo Hyung
Oh, Sangeun
Shin, Insik
Lee, Sunjae
author_facet Im, Youngmin
Jo, Byeongung
Wi, Jaeyoung
Baek, Seungwoo
Min, Tae Hoon
Lee, Joo Hyung
Oh, Sangeun
Shin, Insik
Lee, Sunjae
contents Mobile GUI Agents, AI agents capable of interacting with mobile applications on behalf of users, have the potential to transform human computer interaction. However, current evaluation practices for GUI agents face two fundamental limitations. First, they either rely on single path offline benchmarks or online live benchmarks. Offline benchmarks using static, single path annotated datasets unfairly penalize valid alternative actions, while online benchmarks suffer from poor scalability and reproducibility due to the dynamic and unpredictable nature of live evaluation. Second, existing benchmarks treat agents as monolithic black boxes, overlooking the contributions of individual components, which often leads to unfair comparisons or obscures key performance bottlenecks. To address these limitations, we present MobiBench, the first modular and multi path aware offline benchmarking framework for mobile GUI agents that enables high fidelity, scalable, and reproducible evaluation entirely in offline settings. Our experiments demonstrate that MobiBench achieves 94.72 percent agreement with human evaluators, on par with carefully engineered online benchmarks, while preserving the scalability and reproducibility of static offline benchmarks. Furthermore, our comprehensive module level analysis uncovers several key insights, including a systematic evaluation of diverse techniques used in mobile GUI agents, optimal module configurations across model scales, the inherent limitations of current LFMs, and actionable guidelines for designing more capable and cost efficient mobile agents.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12634
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
Im, Youngmin
Jo, Byeongung
Wi, Jaeyoung
Baek, Seungwoo
Min, Tae Hoon
Lee, Joo Hyung
Oh, Sangeun
Shin, Insik
Lee, Sunjae
Artificial Intelligence
Mobile GUI Agents, AI agents capable of interacting with mobile applications on behalf of users, have the potential to transform human computer interaction. However, current evaluation practices for GUI agents face two fundamental limitations. First, they either rely on single path offline benchmarks or online live benchmarks. Offline benchmarks using static, single path annotated datasets unfairly penalize valid alternative actions, while online benchmarks suffer from poor scalability and reproducibility due to the dynamic and unpredictable nature of live evaluation. Second, existing benchmarks treat agents as monolithic black boxes, overlooking the contributions of individual components, which often leads to unfair comparisons or obscures key performance bottlenecks. To address these limitations, we present MobiBench, the first modular and multi path aware offline benchmarking framework for mobile GUI agents that enables high fidelity, scalable, and reproducible evaluation entirely in offline settings. Our experiments demonstrate that MobiBench achieves 94.72 percent agreement with human evaluators, on par with carefully engineered online benchmarks, while preserving the scalability and reproducibility of static offline benchmarks. Furthermore, our comprehensive module level analysis uncovers several key insights, including a systematic evaluation of diverse techniques used in mobile GUI agents, optimal module configurations across model scales, the inherent limitations of current LFMs, and actionable guidelines for designing more capable and cost efficient mobile agents.
title MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2512.12634