MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Luyuan, Deng, Yongyu, Zha, Yiwei, Mao, Guodong, Wang, Qinmin, Min, Tianchen, Chen, Wei, Chen, Shoufa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
von: Gong, Yichen, et al.
Veröffentlicht: (2026)
von: Gong, Yichen, et al.
Veröffentlicht: (2026)
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
von: Liu, Guangyi, et al.
Veröffentlicht: (2025)
von: Liu, Guangyi, et al.
Veröffentlicht: (2025)
Be Friendly, Not Friends: How LLM Sycophancy Shapes User Trust
von: Sun, Yuan, et al.
Veröffentlicht: (2025)
von: Sun, Yuan, et al.
Veröffentlicht: (2025)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
von: Sun, Jiazheng, et al.
Veröffentlicht: (2026)
von: Sun, Jiazheng, et al.
Veröffentlicht: (2026)
AppAgent v2: Advanced Agent for Flexible Mobile Interactions
von: Li, Yanda, et al.
Veröffentlicht: (2024)
von: Li, Yanda, et al.
Veröffentlicht: (2024)
FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
von: Yang, Qinglong, et al.
Veröffentlicht: (2025)
von: Yang, Qinglong, et al.
Veröffentlicht: (2025)
Benchmarking Mobile Device Control Agents across Diverse Configurations
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
von: Wu, Liangxuan, et al.
Veröffentlicht: (2025)
von: Wu, Liangxuan, et al.
Veröffentlicht: (2025)
From Human Negotiation to Agent Negotiation: Personal Mobility Agents in Automated Traffic
von: Jansen, Pascal
Veröffentlicht: (2026)
von: Jansen, Pascal
Veröffentlicht: (2026)
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
von: Tang, Yuanbo, et al.
Veröffentlicht: (2026)
von: Tang, Yuanbo, et al.
Veröffentlicht: (2026)
AppGen: Mobility-aware App Usage Behavior Generation for Mobile Users
von: Huang, Zihan, et al.
Veröffentlicht: (2024)
von: Huang, Zihan, et al.
Veröffentlicht: (2024)
Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLM
von: Liu, Zhe, et al.
Veröffentlicht: (2024)
von: Liu, Zhe, et al.
Veröffentlicht: (2024)
To Embody or Not: The Effect Of Embodiment On User Perception Of LLM-based Conversational Agents
von: Wang, Kyra, et al.
Veröffentlicht: (2025)
von: Wang, Kyra, et al.
Veröffentlicht: (2025)
TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
von: Lu, Yuheng, et al.
Veröffentlicht: (2025)
von: Lu, Yuheng, et al.
Veröffentlicht: (2025)
AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2026)
MobileExperts: A Dynamic Tool-Enabled Agent Team in Mobile Devices
von: Zhang, Jiayi, et al.
Veröffentlicht: (2024)
von: Zhang, Jiayi, et al.
Veröffentlicht: (2024)
SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
von: Guo, Longjie, et al.
Veröffentlicht: (2025)
von: Guo, Longjie, et al.
Veröffentlicht: (2025)
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration
von: Pan, Bo, et al.
Veröffentlicht: (2024)
von: Pan, Bo, et al.
Veröffentlicht: (2024)
MobileViews: A Million-scale and Diverse Mobile GUI Dataset
von: Gao, Longxi, et al.
Veröffentlicht: (2024)
von: Gao, Longxi, et al.
Veröffentlicht: (2024)
Improving Collaborative Filtering Recommendation via Graph Learning
von: Wang, Yongyu
Veröffentlicht: (2023)
von: Wang, Yongyu
Veröffentlicht: (2023)
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
von: Ai, Kuangshi, et al.
Veröffentlicht: (2026)
von: Ai, Kuangshi, et al.
Veröffentlicht: (2026)
Free Lunch for User Experience: Crowdsourcing Agents for Scalable User Studies
von: Liu, Siyang, et al.
Veröffentlicht: (2025)
von: Liu, Siyang, et al.
Veröffentlicht: (2025)
Human-Centered LLM-Agent User Interface: A Position Paper
von: Chin, Daniel, et al.
Veröffentlicht: (2024)
von: Chin, Daniel, et al.
Veröffentlicht: (2024)
Multi-User Mobile Augmented Reality for Cardiovascular Surgical Planning
von: Mehta, Pratham, et al.
Veröffentlicht: (2024)
von: Mehta, Pratham, et al.
Veröffentlicht: (2024)
GTA: Generative Traffic Agents for Simulating Realistic Mobility Behavior
von: Lämmer, Simon, et al.
Veröffentlicht: (2026)
von: Lämmer, Simon, et al.
Veröffentlicht: (2026)
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
von: Fan, Jingru, et al.
Veröffentlicht: (2025)
von: Fan, Jingru, et al.
Veröffentlicht: (2025)
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
von: Zhu, Ming, et al.
Veröffentlicht: (2026)
von: Zhu, Ming, et al.
Veröffentlicht: (2026)
User Understanding of Privacy Permissions in Mobile Augmented Reality: Perceptions and Misconceptions
von: Paneva, Viktorija, et al.
Veröffentlicht: (2025)
von: Paneva, Viktorija, et al.
Veröffentlicht: (2025)
Influence of Interactivity in Shaping User Experience and Social Acceptance of Mobile XR
von: Kojić, Tanja, et al.
Veröffentlicht: (2026)
von: Kojić, Tanja, et al.
Veröffentlicht: (2026)
The Influence of UX Design on User Retention and Conversion Rates in Mobile Apps
von: Majumder, Aaditya Shankar
Veröffentlicht: (2025)
von: Majumder, Aaditya Shankar
Veröffentlicht: (2025)
AgentLens: Visual Analysis for Agent Behaviors in LLM-based Autonomous Systems
von: Lu, Jiaying, et al.
Veröffentlicht: (2024)
von: Lu, Jiaying, et al.
Veröffentlicht: (2024)
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2026)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2026)
Tap-to-Adapt: Learning User-Aligned Response Timing for Speech Agents
von: He, Zihong, et al.
Veröffentlicht: (2026)
von: He, Zihong, et al.
Veröffentlicht: (2026)
How Personal Characteristics Shape User Exploration of Diverse Movie Recommendations with a LLM-Based Multi-Agent System
von: Zhou, Yufan, et al.
Veröffentlicht: (2026)
von: Zhou, Yufan, et al.
Veröffentlicht: (2026)
PrefIx: Understand and Adapt to User Preference in Human-Agent Interaction
von: Li, Jialin, et al.
Veröffentlicht: (2026)
von: Li, Jialin, et al.
Veröffentlicht: (2026)
ReorderBench: A Benchmark for Matrix Reordering
von: Zhu, Jiangning, et al.
Veröffentlicht: (2024)
von: Zhu, Jiangning, et al.
Veröffentlicht: (2024)
Listen to the Voices of Everyday Users: Democratizing Privacy Ratings for Sensitive Data Access in Mobile Apps
von: Wang, Liu, et al.
Veröffentlicht: (2026)
von: Wang, Liu, et al.
Veröffentlicht: (2026)
AI-Gadget Kit: Integrating Swarm User Interfaces with LLM-driven Agents for Rich Tabletop Game Applications
von: Guo, Yijie, et al.
Veröffentlicht: (2024)
von: Guo, Yijie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
von: Gong, Yichen, et al.
Veröffentlicht: (2026) -
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
von: Liu, Guangyi, et al.
Veröffentlicht: (2025) -
Be Friendly, Not Friends: How LLM Sycophancy Shapes User Trust
von: Sun, Yuan, et al.
Veröffentlicht: (2025) -
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
von: Sun, Jiazheng, et al.
Veröffentlicht: (2026) -
AppAgent v2: Advanced Agent for Flexible Mobile Interactions
von: Li, Yanda, et al.
Veröffentlicht: (2024)