GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Chenrui, Yu, Zedong, Gao, Zhi, Feng, Ruining, Liu, Enqi, Wu, Yuwei, Jia, Yunde, Xiang, Liuyu, He, Zhaofeng, Li, Qing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
von: Liu, Enqi, et al.
Veröffentlicht: (2026)
von: Liu, Enqi, et al.
Veröffentlicht: (2026)
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
von: Xie, Rui, et al.
Veröffentlicht: (2026)
von: Xie, Rui, et al.
Veröffentlicht: (2026)
GUI-PRA: Process Reward Agent for GUI Tasks
von: Xiong, Tao, et al.
Veröffentlicht: (2025)
von: Xiong, Tao, et al.
Veröffentlicht: (2025)
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge
von: Du, Yuntao, et al.
Veröffentlicht: (2025)
von: Du, Yuntao, et al.
Veröffentlicht: (2025)
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
von: Liu, Guangyi, et al.
Veröffentlicht: (2026)
von: Liu, Guangyi, et al.
Veröffentlicht: (2026)
GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
von: Xie, Bin, et al.
Veröffentlicht: (2025)
von: Xie, Bin, et al.
Veröffentlicht: (2025)
Memory-Centric Embodied Question Answering
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025)
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025)
Long-Horizon Visual Imitation Learning via Plan and Code Reflection
von: Chen, Quan, et al.
Veröffentlicht: (2025)
von: Chen, Quan, et al.
Veröffentlicht: (2025)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
von: Nong, Songqin, et al.
Veröffentlicht: (2025)
von: Nong, Songqin, et al.
Veröffentlicht: (2025)
Large-Scale Riemannian Meta-Optimization via Subspace Adaptation
von: Yu, Peilin, et al.
Veröffentlicht: (2025)
von: Yu, Peilin, et al.
Veröffentlicht: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
GraphPilot: GUI Task Automation with One-Step LLM Reasoning Powered by Knowledge Graph
von: Yu, Mingxian, et al.
Veröffentlicht: (2026)
von: Yu, Mingxian, et al.
Veröffentlicht: (2026)
LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning
von: Wu, Yubin, et al.
Veröffentlicht: (2026)
von: Wu, Yubin, et al.
Veröffentlicht: (2026)
TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks
von: Yin, Yiwen, et al.
Veröffentlicht: (2025)
von: Yin, Yiwen, et al.
Veröffentlicht: (2025)
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
von: Wang, Yangyue, et al.
Veröffentlicht: (2026)
von: Wang, Yangyue, et al.
Veröffentlicht: (2026)
What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs
von: Lin, Jiaping, et al.
Veröffentlicht: (2026)
von: Lin, Jiaping, et al.
Veröffentlicht: (2026)
GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
von: Liu, Tao, et al.
Veröffentlicht: (2025)
von: Liu, Tao, et al.
Veröffentlicht: (2025)
Curvature Learning for Generalization of Hyperbolic Neural Networks
von: Fan, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Fan, Xiaomeng, et al.
Veröffentlicht: (2025)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
von: Zhang, Xintong, et al.
Veröffentlicht: (2025)
von: Zhang, Xintong, et al.
Veröffentlicht: (2025)
GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents
von: Wang, Yanxi, et al.
Veröffentlicht: (2026)
von: Wang, Yanxi, et al.
Veröffentlicht: (2026)
POINTS-GUI-G: GUI-Grounding Journey
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
von: Zhao, Zhongyin, et al.
Veröffentlicht: (2026)
GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
von: Chen, Cong, et al.
Veröffentlicht: (2025)
von: Chen, Cong, et al.
Veröffentlicht: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Hyperbolic Dual Feature Augmentation for Open-Environment
von: Yu, Peilin, et al.
Veröffentlicht: (2025)
von: Yu, Peilin, et al.
Veröffentlicht: (2025)
GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control
von: Wei, Jingxuan, et al.
Veröffentlicht: (2026)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2026)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
von: Tang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Tang, Jiaqi, et al.
Veröffentlicht: (2025)
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
von: Li, Hongxin, et al.
Veröffentlicht: (2025)
Rethinking Class-Incremental Learning from a Dynamic Imbalanced Learning Perspective
von: Wang, Leyuan, et al.
Veröffentlicht: (2024)
von: Wang, Leyuan, et al.
Veröffentlicht: (2024)
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
von: Li, Pengxiang, et al.
Veröffentlicht: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents
von: Chai, Yuxiang, et al.
Veröffentlicht: (2026)
von: Chai, Yuxiang, et al.
Veröffentlicht: (2026)
GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
von: Cao, Yuan, et al.
Veröffentlicht: (2026)
von: Cao, Yuan, et al.
Veröffentlicht: (2026)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
von: Liu, Yichao, et al.
Veröffentlicht: (2026)
von: Liu, Yichao, et al.
Veröffentlicht: (2026)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
von: Liu, Shuodi, et al.
Veröffentlicht: (2025)
von: Liu, Shuodi, et al.
Veröffentlicht: (2025)
MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution
von: Sun, Libo, et al.
Veröffentlicht: (2026)
von: Sun, Libo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
von: Li, Pengxiang, et al.
Veröffentlicht: (2025) -
Benchmarking and Improving GUI Agents in High-Dynamic Environments
von: Liu, Enqi, et al.
Veröffentlicht: (2026) -
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
von: Xie, Rui, et al.
Veröffentlicht: (2026) -
GUI-PRA: Process Reward Agent for GUI Tasks
von: Xiong, Tao, et al.
Veröffentlicht: (2025) -
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge
von: Du, Yuntao, et al.
Veröffentlicht: (2025)