AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sun, Jiazheng, Li, Mingxuan, Zhang, Yingying, Niu, Jiayang, Wu, Yachen, Jin, Ruihan, Lei, Shuyu, Tan, Pengrongrui, Zhang, Zongyu, Wang, Ruoyi, Yang, Jiachen, Yang, Boyu, Liu, Jiacheng, Peng, Xin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Robust, Observable, and Evolvable Agentic Systems Engineering: A Principled Framework Validated via the Fairy GUI Agent
par: Sun, Jiazheng, et autres
Publié: (2025)
par: Sun, Jiazheng, et autres
Publié: (2025)
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
par: Wang, Pengyu, et autres
Publié: (2026)
par: Wang, Pengyu, et autres
Publié: (2026)
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
par: Liu, Guangyi, et autres
Publié: (2026)
par: Liu, Guangyi, et autres
Publié: (2026)
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
par: Zhu, Jiachen, et autres
Publié: (2026)
par: Zhu, Jiachen, et autres
Publié: (2026)
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
par: Wu, Qinzhuo, et autres
Publié: (2026)
par: Wu, Qinzhuo, et autres
Publié: (2026)
MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
par: Im, Youngmin, et autres
Publié: (2025)
par: Im, Youngmin, et autres
Publié: (2025)
BaiJia: A Large-Scale Role-Playing Agent Corpus of Chinese Historical Characters
par: Bai, Ting, et autres
Publié: (2024)
par: Bai, Ting, et autres
Publié: (2024)
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
par: Zhao, Pengxiang, et autres
Publié: (2025)
par: Zhao, Pengxiang, et autres
Publié: (2025)
Advanced Multi‐Blockchain Methods for Privacy‐Enhanced IoT ‐Driven Supply Chain Collaboration
par: Jiazheng Lin, et autres
Publié: (2026)
par: Jiazheng Lin, et autres
Publié: (2026)
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
par: Li, Yang, et autres
Publié: (2026)
par: Li, Yang, et autres
Publié: (2026)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
par: Zhang, Ziyun, et autres
Publié: (2026)
par: Zhang, Ziyun, et autres
Publié: (2026)
One-Shot Learning as Instruction Data Prospector for Large Language Models
par: Li, Yunshui, et autres
Publié: (2023)
par: Li, Yunshui, et autres
Publié: (2023)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
par: Gong, Yichen, et autres
Publié: (2026)
par: Gong, Yichen, et autres
Publié: (2026)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
par: Shi, Yucheng, et autres
Publié: (2025)
par: Shi, Yucheng, et autres
Publié: (2025)
The cohomology of certain intermediate strata of Kottwitz varieties
par: Liu, Yachen
Publié: (2025)
par: Liu, Yachen
Publié: (2025)
An automorphic description of the zeta function of the basic stratum of certain Kottwitz varieties
par: Liu, Yachen
Publié: (2024)
par: Liu, Yachen
Publié: (2024)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
par: Chai, Yuxiang, et autres
Publié: (2024)
par: Chai, Yuxiang, et autres
Publié: (2024)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
par: Pei, Siqi, et autres
Publié: (2026)
par: Pei, Siqi, et autres
Publié: (2026)
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
par: Sun, Yuchen, et autres
Publié: (2025)
par: Sun, Yuchen, et autres
Publié: (2025)
Aria-UI: Visual Grounding for GUI Instructions
par: Yang, Yuhao, et autres
Publié: (2024)
par: Yang, Yuhao, et autres
Publié: (2024)
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
par: Chen, Ruihan, et autres
Publié: (2025)
par: Chen, Ruihan, et autres
Publié: (2025)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
par: Zhang, Linhao, et autres
Publié: (2025)
par: Zhang, Linhao, et autres
Publié: (2025)
Mobile GUI Agents under Real-world Threats: Are We There Yet?
par: Liu, Guohong, et autres
Publié: (2025)
par: Liu, Guohong, et autres
Publié: (2025)
MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices
par: Jiang, Jianwen, et autres
Publié: (2024)
par: Jiang, Jianwen, et autres
Publié: (2024)
MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
par: Zhao, Ruihan, et autres
Publié: (2025)
par: Zhao, Ruihan, et autres
Publié: (2025)
AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment
par: Ivanova, Anastasiia, et autres
Publié: (2025)
par: Ivanova, Anastasiia, et autres
Publié: (2025)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
par: Xu, Yifan, et autres
Publié: (2025)
par: Xu, Yifan, et autres
Publié: (2025)
Massive Acquisition of Ultraviolet Color Excess Information from GALEX and UVOT Bands
par: Yang, Dongliang, et autres
Publié: (2025)
par: Yang, Dongliang, et autres
Publié: (2025)
Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents
par: Yang, Yulong, et autres
Publié: (2024)
par: Yang, Yulong, et autres
Publié: (2024)
ProBench: Benchmarking GUI Agents with Accurate Process Information
par: Yang, Leyang, et autres
Publié: (2025)
par: Yang, Leyang, et autres
Publié: (2025)
PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent
par: Nie, Hongyi, et autres
Publié: (2026)
par: Nie, Hongyi, et autres
Publié: (2026)
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
par: Shi, Chenrui, et autres
Publié: (2025)
par: Shi, Chenrui, et autres
Publié: (2025)
LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
par: Liu, Guangyi, et autres
Publié: (2025)
par: Liu, Guangyi, et autres
Publié: (2025)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
par: Tang, Liujian, et autres
Publié: (2025)
par: Tang, Liujian, et autres
Publié: (2025)
GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents
par: Wang, Yanxi, et autres
Publié: (2026)
par: Wang, Yanxi, et autres
Publié: (2026)
WildLMa: Long Horizon Loco-Manipulation in the Wild
par: Qiu, Ri-Zhao, et autres
Publié: (2024)
par: Qiu, Ri-Zhao, et autres
Publié: (2024)
GUI Agents for Continual Game Generation
par: Huang, Yixu, et autres
Publié: (2026)
par: Huang, Yixu, et autres
Publié: (2026)
PACKMOL- GUI: An All-in-One VMD Interface for Efficient Molecular Packing
par: Huang, Jian, et autres
Publié: (2024)
par: Huang, Jian, et autres
Publié: (2024)
AmbiSQL: Interactive Ambiguity Detection and Resolution for Text-to-SQL
par: Ding, Zhongjun, et autres
Publié: (2025)
par: Ding, Zhongjun, et autres
Publié: (2025)
Documents similaires
-
Robust, Observable, and Evolvable Agentic Systems Engineering: A Principled Framework Validated via the Fairy GUI Agent
par: Sun, Jiazheng, et autres
Publié: (2025) -
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
par: Wang, Pengyu, et autres
Publié: (2026) -
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
par: Liu, Guangyi, et autres
Publié: (2026) -
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
par: Zhu, Jiachen, et autres
Publié: (2026) -
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
par: Wu, Qinzhuo, et autres
Publié: (2026)