AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Hongxin, Wang, Xiping, Su, Jingran, Ju, Zheng, Chen, Yuntao, Li, Qing, Zhang, Zhaoxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction
von: Li, Hongxin, et al.
Veröffentlicht: (2026)
von: Li, Hongxin, et al.
Veröffentlicht: (2026)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
von: Wang, Xuehui, et al.
Veröffentlicht: (2025)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
von: Liu, Enqi, et al.
Veröffentlicht: (2026)
von: Liu, Enqi, et al.
Veröffentlicht: (2026)
GEBench: Benchmarking Image Generation Models as GUI Environments
von: Li, Haodong, et al.
Veröffentlicht: (2026)
von: Li, Haodong, et al.
Veröffentlicht: (2026)
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
von: Ji, Fengxian, et al.
Veröffentlicht: (2025)
von: Ji, Fengxian, et al.
Veröffentlicht: (2025)
VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
von: Zhou, Beitong, et al.
Veröffentlicht: (2025)
von: Zhou, Beitong, et al.
Veröffentlicht: (2025)
POINTS-GUI-G: GUI-Grounding Journey
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
von: Li, Junlong, et al.
Veröffentlicht: (2026)
von: Li, Junlong, et al.
Veröffentlicht: (2026)
GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
von: Liu, Tao, et al.
Veröffentlicht: (2025)
von: Liu, Tao, et al.
Veröffentlicht: (2025)
Step-GUI Technical Report
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
von: Yan, Haolong, et al.
Veröffentlicht: (2025)
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
von: Yao, Ruilin, et al.
Veröffentlicht: (2026)
von: Yao, Ruilin, et al.
Veröffentlicht: (2026)
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
von: Sun, Yuchen, et al.
Veröffentlicht: (2025)
von: Sun, Yuchen, et al.
Veröffentlicht: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
von: Lu, Quanfeng, et al.
Veröffentlicht: (2024)
von: Lu, Quanfeng, et al.
Veröffentlicht: (2024)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
Continual GUI Agents
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
von: Liu, Ziwei, et al.
Veröffentlicht: (2026)
DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
von: Chen, Yuntao, et al.
Veröffentlicht: (2024)
von: Chen, Yuntao, et al.
Veröffentlicht: (2024)
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
von: Zhao, Haoren, et al.
Veröffentlicht: (2026)
von: Zhao, Haoren, et al.
Veröffentlicht: (2026)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
von: Shao, Rui, et al.
Veröffentlicht: (2026)
von: Shao, Rui, et al.
Veröffentlicht: (2026)
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
von: Fu, Chaoyou, et al.
Veröffentlicht: (2026)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2026)
UItron: Foundational GUI Agent with Advanced Perception and Planning
von: Zeng, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zeng, Zhixiong, et al.
Veröffentlicht: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
von: Bu, Tianpeng, et al.
Veröffentlicht: (2026)
von: Bu, Tianpeng, et al.
Veröffentlicht: (2026)
Auto-scaling Continuous Memory for GUI Agent
von: Wu, Wenyi, et al.
Veröffentlicht: (2025)
von: Wu, Wenyi, et al.
Veröffentlicht: (2025)
FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
von: Ji, Fengxian, et al.
Veröffentlicht: (2026)
von: Ji, Fengxian, et al.
Veröffentlicht: (2026)
HLG: Comprehensive 3D Room Construction via Hierarchical Layout Generation
von: Wang, Xiping, et al.
Veröffentlicht: (2025)
von: Wang, Xiping, et al.
Veröffentlicht: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
von: Li, Hongxin, et al.
Veröffentlicht: (2025) -
UIPro: Unleashing Superior Interaction Capability For GUI Agents
von: Li, Hongxin, et al.
Veröffentlicht: (2025) -
GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction
von: Li, Hongxin, et al.
Veröffentlicht: (2026) -
MP-GUI: Modality Perception with MLLMs for GUI Understanding
von: Wang, Ziwei, et al.
Veröffentlicht: (2025) -
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
von: Li, Yang, et al.
Veröffentlicht: (2026)