GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yang, Liu, Yuchen, Lu, Haoyu, Xia, Zhiqiang, Wang, Hongzhen, Han, Kaiyang, Yang, Changpeng, Wu, Jinyang, Xu, Jiaming, Shi, Runyu, Huang, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
by: Wu, Qinzhuo, et al.
Published: (2026)
by: Wu, Qinzhuo, et al.
Published: (2026)
ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devices
by: Kong, Dezhi, et al.
Published: (2026)
by: Kong, Dezhi, et al.
Published: (2026)
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
by: Liu, Guangyi, et al.
Published: (2026)
by: Liu, Guangyi, et al.
Published: (2026)
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
by: Chen, Ruihan, et al.
Published: (2025)
by: Chen, Ruihan, et al.
Published: (2025)
From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks
by: Yang, Changpeng, et al.
Published: (2025)
by: Yang, Changpeng, et al.
Published: (2025)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
by: Wang, Xuehui, et al.
Published: (2025)
by: Wang, Xuehui, et al.
Published: (2025)
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
by: Sun, Liangtai, et al.
Published: (2022)
by: Sun, Liangtai, et al.
Published: (2022)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
by: Liu, Enqi, et al.
Published: (2026)
by: Liu, Enqi, et al.
Published: (2026)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
by: Shi, Yucheng, et al.
Published: (2025)
by: Shi, Yucheng, et al.
Published: (2025)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)
by: Tang, Liujian, et al.
Published: (2025)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)
by: Henry, Felix, et al.
Published: (2026)
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
by: Zhu, Jiachen, et al.
Published: (2026)
by: Zhu, Jiachen, et al.
Published: (2026)
TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine
by: Huang, Tianai, et al.
Published: (2025)
by: Huang, Tianai, et al.
Published: (2025)
MobileFlow: A Multimodal LLM For Mobile GUI Agent
by: Nong, Songqin, et al.
Published: (2024)
by: Nong, Songqin, et al.
Published: (2024)
MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation
by: Li, Ning, et al.
Published: (2025)
by: Li, Ning, et al.
Published: (2025)
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
by: Sun, Yuchen, et al.
Published: (2025)
by: Sun, Yuchen, et al.
Published: (2025)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
by: Sun, Jiazheng, et al.
Published: (2026)
by: Sun, Jiazheng, et al.
Published: (2026)
GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
by: Yang, Jingqi, et al.
Published: (2025)
by: Yang, Jingqi, et al.
Published: (2025)
GUITester: Enabling GUI Agents for Exploratory Defect Discovery
by: Gao, Yifei, et al.
Published: (2026)
by: Gao, Yifei, et al.
Published: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
GUI-360$^\circ$: A Comprehensive Dataset and Benchmark for Computer-Using Agents
by: Mu, Jian, et al.
Published: (2025)
by: Mu, Jian, et al.
Published: (2025)
Mobile-Agent-v3: Fundamental Agents for GUI Automation
by: Ye, Jiabo, et al.
Published: (2025)
by: Ye, Jiabo, et al.
Published: (2025)
MobileDreamer: Generative Sketch World Model for GUI Agent
by: Cao, Yilin, et al.
Published: (2026)
by: Cao, Yilin, et al.
Published: (2026)
GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
by: Im, Youngmin, et al.
Published: (2025)
by: Im, Youngmin, et al.
Published: (2025)
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation
by: Qu, Heng, et al.
Published: (2026)
by: Qu, Heng, et al.
Published: (2026)
FedGUI: Benchmarking Federated GUI Agents across Heterogeneous Platforms, Devices, and Operating Systems
by: Wang, Wenhao, et al.
Published: (2026)
by: Wang, Wenhao, et al.
Published: (2026)
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
by: Zhao, Pengxiang, et al.
Published: (2025)
by: Zhao, Pengxiang, et al.
Published: (2025)
Faithful Mobile GUI Agents with Guided Advantage Estimator
by: Hu, Haowen, et al.
Published: (2026)
by: Hu, Haowen, et al.
Published: (2026)
GUI-PRA: Process Reward Agent for GUI Tasks
by: Xiong, Tao, et al.
Published: (2025)
by: Xiong, Tao, et al.
Published: (2025)
FuncDroid: Towards Inter-Functional Flows for Comprehensive Mobile App GUI Testing
by: He, Jinlong, et al.
Published: (2026)
by: He, Jinlong, et al.
Published: (2026)
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
by: Li, Hongxin, et al.
Published: (2026)
by: Li, Hongxin, et al.
Published: (2026)
AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
by: Zhang, Zhong, et al.
Published: (2025)
by: Zhang, Zhong, et al.
Published: (2025)
BEAP-Agent: Backtrackable Execution and Adaptive Planning for GUI Agents
by: Lu, Ziyu, et al.
Published: (2026)
by: Lu, Ziyu, et al.
Published: (2026)
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
by: Zhao, Kangjia, et al.
Published: (2024)
by: Zhao, Kangjia, et al.
Published: (2024)
Similar Items
-
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
by: Wu, Qinzhuo, et al.
Published: (2026) -
ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devices
by: Kong, Dezhi, et al.
Published: (2026) -
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
by: Liu, Guangyi, et al.
Published: (2026) -
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
by: Chen, Ruihan, et al.
Published: (2025) -
From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks
by: Yang, Changpeng, et al.
Published: (2025)