OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Henry, Felix, Lin, Xiaochen, Zhu, Jiangyou, Yangfan, Zhang, Bingqian, Chen, Min, Huang, Shiyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
GUI Agents: A Survey
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
API Agents vs. GUI Agents: Divergence and Convergence
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)
by: Tang, Liujian, et al.
Published: (2025)
AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
by: Kim, Jeonghyeon, et al.
Published: (2026)
by: Kim, Jeonghyeon, et al.
Published: (2026)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
GUI Agents with Foundation Models: A Comprehensive Survey
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024)
by: Yang, Yuhao, et al.
Published: (2024)
TinyClick: Single-Turn Agent for Empowering GUI Automation
by: Pawlowski, Pawel, et al.
Published: (2024)
by: Pawlowski, Pawel, et al.
Published: (2024)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
by: Chai, Yuxiang, et al.
Published: (2024)
by: Chai, Yuxiang, et al.
Published: (2024)
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
by: Gong, Yichen, et al.
Published: (2026)
by: Gong, Yichen, et al.
Published: (2026)
Large Language Model-Brained GUI Agents: A Survey
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
by: Sun, Jiazheng, et al.
Published: (2026)
by: Sun, Jiazheng, et al.
Published: (2026)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
by: Tang, Jingyu, et al.
Published: (2025)
by: Tang, Jingyu, et al.
Published: (2025)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
GUI Agents for Continual Game Generation
by: Huang, Yixu, et al.
Published: (2026)
by: Huang, Yixu, et al.
Published: (2026)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
History-Aware Reasoning for GUI Agents
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
Qualitative Evaluation of LLM-Designed GUI
by: Sawicki, Bartosz, et al.
Published: (2026)
by: Sawicki, Bartosz, et al.
Published: (2026)
TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks
by: Yin, Yiwen, et al.
Published: (2025)
by: Yin, Yiwen, et al.
Published: (2025)
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
by: Hu, Haobo, et al.
Published: (2026)
by: Hu, Haobo, et al.
Published: (2026)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
by: Wu, Zongru, et al.
Published: (2025)
by: Wu, Zongru, et al.
Published: (2025)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
by: Lee, Jungjae, et al.
Published: (2025)
by: Lee, Jungjae, et al.
Published: (2025)
Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment
by: Sun, Yuchen, et al.
Published: (2026)
by: Sun, Yuchen, et al.
Published: (2026)
OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering
by: Li, Jiahao Nick, et al.
Published: (2024)
by: Li, Jiahao Nick, et al.
Published: (2024)
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
by: Yang, Saelyne, et al.
Published: (2026)
by: Yang, Saelyne, et al.
Published: (2026)
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
by: Liu, Zhilin, et al.
Published: (2026)
by: Liu, Zhilin, et al.
Published: (2026)
A Survey on (M)LLM-Based GUI Agents
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
by: Liao, Zeyi, et al.
Published: (2025)
by: Liao, Zeyi, et al.
Published: (2025)
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
by: Qin, Yujia, et al.
Published: (2025)
by: Qin, Yujia, et al.
Published: (2025)
ControlGUI: Guiding Generative GUI Exploration through Perceptual Visual Flow
by: Garg, Aryan, et al.
Published: (2025)
by: Garg, Aryan, et al.
Published: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
by: Sun, Qiang, et al.
Published: (2024)
by: Sun, Qiang, et al.
Published: (2024)
GUICourse: From General Vision Language Models to Versatile GUI Agents
by: Chen, Wentong, et al.
Published: (2024)
by: Chen, Wentong, et al.
Published: (2024)
OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs
by: Li, Jiahao Nick, et al.
Published: (2024)
by: Li, Jiahao Nick, et al.
Published: (2024)
Less is More: Empowering GUI Agent with Context-Aware Simplification
by: Chen, Gongwei, et al.
Published: (2025)
by: Chen, Gongwei, et al.
Published: (2025)
Robust, Observable, and Evolvable Agentic Systems Engineering: A Principled Framework Validated via the Fairy GUI Agent
by: Sun, Jiazheng, et al.
Published: (2025)
by: Sun, Jiazheng, et al.
Published: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
Similar Items
-
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025) -
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024) -
GUI Agents: A Survey
by: Nguyen, Dang, et al.
Published: (2024) -
API Agents vs. GUI Agents: Divergence and Convergence
by: Zhang, Chaoyun, et al.
Published: (2025) -
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)