Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumbhar, Shrinidhi, Liao, Haofu, Appalaraju, Srikar, Singh, Kunwar Yashraj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
von: Ye, Xianhang, et al.
Veröffentlicht: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025)
von: Lei, Bin, et al.
Veröffentlicht: (2025)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
von: Sun, Yuchen, et al.
Veröffentlicht: (2025)
von: Sun, Yuchen, et al.
Veröffentlicht: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhou, Yuqi, et al.
Veröffentlicht: (2025)
ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding
von: Hsieh, ZongHan, et al.
Veröffentlicht: (2025)
von: Hsieh, ZongHan, et al.
Veröffentlicht: (2025)
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants
von: Hu, Junan, et al.
Veröffentlicht: (2026)
von: Hu, Junan, et al.
Veröffentlicht: (2026)
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
von: Rao, Varun Nagaraj, et al.
Veröffentlicht: (2024)
von: Rao, Varun Nagaraj, et al.
Veröffentlicht: (2024)
Visual Test-time Scaling for GUI Agent Grounding
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
von: Luo, Tiange, et al.
Veröffentlicht: (2025)
GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
von: Wu, Penghao, et al.
Veröffentlicht: (2025)
GUICourse: From General Vision Language Models to Versatile GUI Agents
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
von: Chen, Wentong, et al.
Veröffentlicht: (2024)
BAMI: Training-Free Bias Mitigation in GUI Grounding
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
von: Zhang, Borui, et al.
Veröffentlicht: (2026)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
Improved GUI Grounding via Iterative Narrowing
von: Nguyen, Anthony
Veröffentlicht: (2024)
von: Nguyen, Anthony
Veröffentlicht: (2024)
GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
von: Liu, Tao, et al.
Veröffentlicht: (2025)
von: Liu, Tao, et al.
Veröffentlicht: (2025)
TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
OmniParser for Pure Vision Based GUI Agent
von: Lu, Yadong, et al.
Veröffentlicht: (2024)
von: Lu, Yadong, et al.
Veröffentlicht: (2024)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
von: Fan, Yue, et al.
Veröffentlicht: (2025)
von: Fan, Yue, et al.
Veröffentlicht: (2025)
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
von: Ma, Longhui, et al.
Veröffentlicht: (2026)
von: Ma, Longhui, et al.
Veröffentlicht: (2026)
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
von: Chen, Liangyu, et al.
Veröffentlicht: (2025)
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
von: Tang, Fei, et al.
Veröffentlicht: (2026)
von: Tang, Fei, et al.
Veröffentlicht: (2026)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2026)
von: Tang, Fei, et al.
Veröffentlicht: (2026)
Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation
von: Gao, Xinzge, et al.
Veröffentlicht: (2025)
von: Gao, Xinzge, et al.
Veröffentlicht: (2025)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2025)
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
von: Lian, Shuquan, et al.
Veröffentlicht: (2025)
von: Lian, Shuquan, et al.
Veröffentlicht: (2025)
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
von: Wang, Wanfu, et al.
Veröffentlicht: (2025)
von: Wang, Wanfu, et al.
Veröffentlicht: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
von: Xu, Zhou, et al.
Veröffentlicht: (2026)
von: Xu, Zhou, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
von: Park, Joonhyung, et al.
Veröffentlicht: (2025) -
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
von: Ye, Xianhang, et al.
Veröffentlicht: (2025) -
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025) -
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
von: Lei, Bin, et al.
Veröffentlicht: (2025) -
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
von: Pei, Siqi, et al.
Veröffentlicht: (2026)