OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Davydova, Mariya, Jeffries, Daniel, Barker, Patrick, Flores, Arturo Márquez, Ryan, Sinéad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
by: Yang, Xiao, et al.
Published: (2025)
by: Yang, Xiao, et al.
Published: (2025)
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
by: Chen, Ruihan, et al.
Published: (2025)
by: Chen, Ruihan, et al.
Published: (2025)
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
by: Zhao, Kangjia, et al.
Published: (2024)
by: Zhao, Kangjia, et al.
Published: (2024)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)
by: Henry, Felix, et al.
Published: (2026)
See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch
by: Zhang, Xingyi, et al.
Published: (2026)
by: Zhang, Xingyi, et al.
Published: (2026)
ProBench: Benchmarking GUI Agents with Accurate Process Information
by: Yang, Leyang, et al.
Published: (2025)
by: Yang, Leyang, et al.
Published: (2025)
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent
by: Nie, Hongyi, et al.
Published: (2026)
by: Nie, Hongyi, et al.
Published: (2026)
Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
by: Xie, Yuquan, et al.
Published: (2025)
by: Xie, Yuquan, et al.
Published: (2025)
MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
by: Im, Youngmin, et al.
Published: (2025)
by: Im, Youngmin, et al.
Published: (2025)
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation
by: Qu, Heng, et al.
Published: (2026)
by: Qu, Heng, et al.
Published: (2026)
GUI-PRA: Process Reward Agent for GUI Tasks
by: Xiong, Tao, et al.
Published: (2025)
by: Xiong, Tao, et al.
Published: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
GUI-360$^\circ$: A Comprehensive Dataset and Benchmark for Computer-Using Agents
by: Mu, Jian, et al.
Published: (2025)
by: Mu, Jian, et al.
Published: (2025)
EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
by: Mo, Ying, et al.
Published: (2026)
by: Mo, Ying, et al.
Published: (2026)
GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
Anticipatory Planning for Multimodal AI Agents
by: Liang, Yongyuan, et al.
Published: (2026)
by: Liang, Yongyuan, et al.
Published: (2026)
GUI Agents: A Survey
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents
by: Yan, Yunhe, et al.
Published: (2025)
by: Yan, Yunhe, et al.
Published: (2025)
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
by: Zhao, Pengxiang, et al.
Published: (2025)
by: Zhao, Pengxiang, et al.
Published: (2025)
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
by: Zhu, Jiachen, et al.
Published: (2026)
by: Zhu, Jiachen, et al.
Published: (2026)
GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
Executable Agentic Memory for GUI Agent
by: Qin, Zerui, et al.
Published: (2026)
by: Qin, Zerui, et al.
Published: (2026)
PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents
by: Chai, Yuxiang, et al.
Published: (2026)
by: Chai, Yuxiang, et al.
Published: (2026)
LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning
by: Wu, Yubin, et al.
Published: (2026)
by: Wu, Yubin, et al.
Published: (2026)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
by: Liu, Guohong, et al.
Published: (2026)
by: Liu, Guohong, et al.
Published: (2026)
GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
by: Yang, Jingqi, et al.
Published: (2025)
by: Yang, Jingqi, et al.
Published: (2025)
GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
by: Gao, Longxi, et al.
Published: (2025)
by: Gao, Longxi, et al.
Published: (2025)
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
by: Sun, Liangtai, et al.
Published: (2022)
by: Sun, Liangtai, et al.
Published: (2022)
Mobile-Agent-v3: Fundamental Agents for GUI Automation
by: Ye, Jiabo, et al.
Published: (2025)
by: Ye, Jiabo, et al.
Published: (2025)
BEAP-Agent: Backtrackable Execution and Adaptive Planning for GUI Agents
by: Lu, Ziyu, et al.
Published: (2026)
by: Lu, Ziyu, et al.
Published: (2026)
AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
by: Leung, Ho Fai, et al.
Published: (2025)
by: Leung, Ho Fai, et al.
Published: (2025)
Reflection-Based Memory For Web navigation Agents
by: Azam, Ruhana, et al.
Published: (2025)
by: Azam, Ruhana, et al.
Published: (2025)
EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
by: Li, Runze, et al.
Published: (2025)
by: Li, Runze, et al.
Published: (2025)
MagicGUI-RMS: A Multi-Agent Reward Model System for Self-Evolving GUI Agents via Automated Feedback Reflux
by: Li, Zecheng, et al.
Published: (2026)
by: Li, Zecheng, et al.
Published: (2026)
AppAgentX: Evolving GUI Agents as Proficient Smartphone Users
by: Jiang, Wenjia, et al.
Published: (2025)
by: Jiang, Wenjia, et al.
Published: (2025)
GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Similar Items
-
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
by: Yang, Xiao, et al.
Published: (2025) -
MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
by: Chen, Ruihan, et al.
Published: (2025) -
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
by: Zhao, Kangjia, et al.
Published: (2024) -
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026) -
See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch
by: Zhang, Xingyi, et al.
Published: (2026)