Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Zeyi, Lu, Yadong, Gou, Boyu, Sun, Huan, Awadallah, Ahmed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents
by: Pahuja, Vardaan, et al.
Published: (2025)
by: Pahuja, Vardaan, et al.
Published: (2025)
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024)
by: Cheng, Kanzhi, et al.
Published: (2024)
TinyClick: Single-Turn Agent for Empowering GUI Automation
by: Pawlowski, Pawel, et al.
Published: (2024)
by: Pawlowski, Pawel, et al.
Published: (2024)
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024)
by: Yang, Yuhao, et al.
Published: (2024)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
by: Sun, Jiazheng, et al.
Published: (2026)
by: Sun, Jiazheng, et al.
Published: (2026)
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
by: Nong, Songqin, et al.
Published: (2025)
by: Nong, Songqin, et al.
Published: (2025)
Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment
by: Sun, Yuchen, et al.
Published: (2026)
by: Sun, Yuchen, et al.
Published: (2026)
GUI Agents: A Survey
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
by: Henry, Felix, et al.
Published: (2026)
by: Henry, Felix, et al.
Published: (2026)
Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation
by: Hao, Pan, et al.
Published: (2026)
by: Hao, Pan, et al.
Published: (2026)
Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding
by: Tan, Wee Joe, et al.
Published: (2026)
by: Tan, Wee Joe, et al.
Published: (2026)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
by: Tang, Liujian, et al.
Published: (2025)
by: Tang, Liujian, et al.
Published: (2025)
StepWrite: Adaptive Planning for Speech-Driven Text Generation
by: Alaoui, Hamza El, et al.
Published: (2025)
by: Alaoui, Hamza El, et al.
Published: (2025)
An Active Inference Model of Mouse Point-and-Click Behaviour
by: Klar, Markus, et al.
Published: (2025)
by: Klar, Markus, et al.
Published: (2025)
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
by: Zhu, Ming, et al.
Published: (2026)
by: Zhu, Ming, et al.
Published: (2026)
GUI Agents with Foundation Models: A Comprehensive Survey
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
API Agents vs. GUI Agents: Divergence and Convergence
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
Robust, Observable, and Evolvable Agentic Systems Engineering: A Principled Framework Validated via the Fairy GUI Agent
by: Sun, Jiazheng, et al.
Published: (2025)
by: Sun, Jiazheng, et al.
Published: (2025)
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents
by: Cheng, Pengzhou, et al.
Published: (2025)
by: Cheng, Pengzhou, et al.
Published: (2025)
TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks
by: Yin, Yiwen, et al.
Published: (2025)
by: Yin, Yiwen, et al.
Published: (2025)
WinClick: GUI Grounding with Multimodal Large Language Models
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
Steps Towards Satisficing Distributed Dynamic Team Trust
by: Hunt, Edmund R., et al.
Published: (2023)
by: Hunt, Edmund R., et al.
Published: (2023)
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment
by: Li, Chenliang, et al.
Published: (2024)
by: Li, Chenliang, et al.
Published: (2024)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Beyond Value Elicitation: Towards Moral Profiles in Early Requirements Engineering via Role-Playing Games and Anthropologist LLMs
by: De Ninno, Gianluca, et al.
Published: (2025)
by: De Ninno, Gianluca, et al.
Published: (2025)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
by: Tang, Jingyu, et al.
Published: (2025)
by: Tang, Jingyu, et al.
Published: (2025)
VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics
by: Gong, Yichen, et al.
Published: (2026)
by: Gong, Yichen, et al.
Published: (2026)
Qualitative Evaluation of LLM-Designed GUI
by: Sawicki, Bartosz, et al.
Published: (2026)
by: Sawicki, Bartosz, et al.
Published: (2026)
Human-Agent Coordination in Games under Incomplete Information via Multi-Step Intent
by: Chen, Shenghui, et al.
Published: (2024)
by: Chen, Shenghui, et al.
Published: (2024)
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
by: Chai, Yuxiang, et al.
Published: (2024)
by: Chai, Yuxiang, et al.
Published: (2024)
Beyond Autocomplete: Designing CopilotLens Towards Transparent and Explainable AI Coding Agents
by: Ye, Runlong, et al.
Published: (2025)
by: Ye, Runlong, et al.
Published: (2025)
Toward Natural and Companionable Virtual Agents via Cross-Temporal Emotional Modeling
by: Qin, Feier, et al.
Published: (2026)
by: Qin, Feier, et al.
Published: (2026)
Beyond Recommendations: From Backward to Forward AI Support of Pilots' Decision-Making Process
by: Zhang, Zelun Tony, et al.
Published: (2024)
by: Zhang, Zelun Tony, et al.
Published: (2024)
AdaptoML-UX: An Adaptive User-centered GUI-based AutoML Toolkit for Non-AI Experts and HCI Researchers
by: Gomaa, Amr, et al.
Published: (2024)
by: Gomaa, Amr, et al.
Published: (2024)
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
by: Lee, Jungjae, et al.
Published: (2025)
by: Lee, Jungjae, et al.
Published: (2025)
Drag or Traction: Understanding How Designers Appropriate Friction in AI Ideation Outputs
by: Kocaballi, A. Baki, et al.
Published: (2026)
by: Kocaballi, A. Baki, et al.
Published: (2026)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
Similar Items
-
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents
by: Pahuja, Vardaan, et al.
Published: (2025) -
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
by: Cheng, Kanzhi, et al.
Published: (2024) -
TinyClick: Single-Turn Agent for Empowering GUI Automation
by: Pawlowski, Pawel, et al.
Published: (2024) -
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024) -
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
by: Liu, Yuhang, et al.
Published: (2025)