AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Jiahui, Hua, Zhichao, Xia, Yubin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
by: Zhou, Zhiyuan, et al.
Published: (2025)
by: Zhou, Zhiyuan, et al.
Published: (2025)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
by: Boyeau, Pierre, et al.
Published: (2024)
by: Boyeau, Pierre, et al.
Published: (2024)
Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
by: Park, Sangwoo, et al.
Published: (2025)
by: Park, Sangwoo, et al.
Published: (2025)
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
by: Ye, Bowen, et al.
Published: (2026)
by: Ye, Bowen, et al.
Published: (2026)
Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation
by: Bhonsle, Roshita, et al.
Published: (2025)
by: Bhonsle, Roshita, et al.
Published: (2025)
Trustworthy Evaluation of Robotic Manipulation: A New Benchmark and AutoEval Methods
by: Liu, Mengyuan, et al.
Published: (2026)
by: Liu, Mengyuan, et al.
Published: (2026)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
by: Wu, JiaRu, et al.
Published: (2025)
by: Wu, JiaRu, et al.
Published: (2025)
AutoHealth: An Uncertainty-Aware Multi-Agent System for Autonomous Health Data Modeling
by: Xia, Tong, et al.
Published: (2026)
by: Xia, Tong, et al.
Published: (2026)
AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
by: Li, Ziming, et al.
Published: (2024)
by: Li, Ziming, et al.
Published: (2024)
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
by: Chen, Weiyi, et al.
Published: (2026)
by: Chen, Weiyi, et al.
Published: (2026)
AutoAgents: A Framework for Automatic Agent Generation
by: Chen, Guangyao, et al.
Published: (2023)
by: Chen, Guangyao, et al.
Published: (2023)
AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning
by: Fu, Suzhong, et al.
Published: (2026)
by: Fu, Suzhong, et al.
Published: (2026)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
by: Zhang, Mengyuan, et al.
Published: (2024)
by: Zhang, Mengyuan, et al.
Published: (2024)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
by: Liu, Tianjian, et al.
Published: (2025)
by: Liu, Tianjian, et al.
Published: (2025)
DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation
by: Xie, Qianqian, et al.
Published: (2026)
by: Xie, Qianqian, et al.
Published: (2026)
AutoGRAMS: Autonomous Graphical Agent Modeling Software
by: Krause, Ben, et al.
Published: (2024)
by: Krause, Ben, et al.
Published: (2024)
MobiAgent: A Systematic Framework for Customizable Mobile Agents
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
by: Mi, Hongze, et al.
Published: (2025)
by: Mi, Hongze, et al.
Published: (2025)
AgentEval: Generative Agents as Reliable Proxies for Human Evaluation of AI-Generated Content
by: Vu, Thanh, et al.
Published: (2025)
by: Vu, Thanh, et al.
Published: (2025)
AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing
by: Du, Jianda, et al.
Published: (2026)
by: Du, Jianda, et al.
Published: (2026)
AutoPKG: An Automated Framework for Dynamic E-commerce Product-Attribute Knowledge Graph Construction
by: Hongwimol, Pollawat, et al.
Published: (2026)
by: Hongwimol, Pollawat, et al.
Published: (2026)
AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents
by: Henke, Julius
Published: (2025)
by: Henke, Julius
Published: (2025)
SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
by: Jiang, Sihang, et al.
Published: (2026)
by: Jiang, Sihang, et al.
Published: (2026)
Aime: Towards Fully-Autonomous Multi-Agent Framework
by: Shi, Yexuan, et al.
Published: (2025)
by: Shi, Yexuan, et al.
Published: (2025)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
by: Garg, Madhav Krishan, et al.
Published: (2025)
by: Garg, Madhav Krishan, et al.
Published: (2025)
A Concurrent Modular Agent: Framework for Autonomous LLM Agents
by: Maruyama, Norihiro, et al.
Published: (2025)
by: Maruyama, Norihiro, et al.
Published: (2025)
AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
by: Lu, Qinghua, et al.
Published: (2025)
by: Lu, Qinghua, et al.
Published: (2025)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
by: Feng, Yunfei, et al.
Published: (2026)
by: Feng, Yunfei, et al.
Published: (2026)
AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation
by: Panapitiya, Gihan, et al.
Published: (2025)
by: Panapitiya, Gihan, et al.
Published: (2025)
VoiceAgentEval: A Dual-Dimensional Benchmark for Expert-Level Intelligent Voice-Agent Evaluation of Xbench's Professional-Aligned Series
by: Xu, Pengyu, et al.
Published: (2025)
by: Xu, Pengyu, et al.
Published: (2025)
AutoGLM: Autonomous Foundation Agents for GUIs
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
by: Deng, Shihan, et al.
Published: (2024)
by: Deng, Shihan, et al.
Published: (2024)
EvalCards: A Framework for Standardized Evaluation Reporting
by: Dhar, Ruchira, et al.
Published: (2025)
by: Dhar, Ruchira, et al.
Published: (2025)
Autonomous Evaluation and Refinement of Digital Agents
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
DevEval: Evaluating Code Generation in Practical Software Projects
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Evaluating Collaborative and Autonomous Agents in Data-Stream-Supported Coordination of Mobile Crowdsourcing
by: Bruns, Ralf, et al.
Published: (2024)
by: Bruns, Ralf, et al.
Published: (2024)
From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent
by: Shen, Minjie, et al.
Published: (2025)
by: Shen, Minjie, et al.
Published: (2025)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
by: Yu, Linhao, et al.
Published: (2024)
by: Yu, Linhao, et al.
Published: (2024)
AutoAgent: A Fully-Automated and Zero-Code Framework for LLM Agents
by: Tang, Jiabin, et al.
Published: (2025)
by: Tang, Jiabin, et al.
Published: (2025)
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
Similar Items
-
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
by: Zhou, Zhiyuan, et al.
Published: (2025) -
AutoEval Done Right: Using Synthetic Data for Model Evaluation
by: Boyeau, Pierre, et al.
Published: (2024) -
Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
by: Park, Sangwoo, et al.
Published: (2025) -
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
by: Ye, Bowen, et al.
Published: (2026) -
Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation
by: Bhonsle, Roshita, et al.
Published: (2025)