GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongxin, Chen, Yuntao, Zhang, Zhaoxiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
by: Li, Hongxin, et al.
Published: (2025)
by: Li, Hongxin, et al.
Published: (2025)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
by: Li, Hongxin, et al.
Published: (2025)
by: Li, Hongxin, et al.
Published: (2025)
AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
by: Li, Hongxin, et al.
Published: (2026)
by: Li, Hongxin, et al.
Published: (2026)
DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
by: Chen, Yuntao, et al.
Published: (2024)
by: Chen, Yuntao, et al.
Published: (2024)
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)
by: Li, Yingyan, et al.
Published: (2024)
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
by: Shang, Shuyao, et al.
Published: (2025)
by: Shang, Shuyao, et al.
Published: (2025)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
by: Jiang, Zhiyuan, et al.
Published: (2025)
by: Jiang, Zhiyuan, et al.
Published: (2025)
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
by: Zhang, Yan, et al.
Published: (2026)
by: Zhang, Yan, et al.
Published: (2026)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
Weakly Supervised 3D Object Detection with Multi-Stage Generalization
by: He, Jiawei, et al.
Published: (2023)
by: He, Jiawei, et al.
Published: (2023)
Monocular Occupancy Prediction for Scalable Indoor Scenes
by: Yu, Hongxiao, et al.
Published: (2024)
by: Yu, Hongxiao, et al.
Published: (2024)
ClickAttention: Click Region Similarity Guided Interactive Segmentation
by: Xu, Long, et al.
Published: (2024)
by: Xu, Long, et al.
Published: (2024)
POINTS-GUI-G: GUI-Grounding Journey
by: Zhao, Zhongyin, et al.
Published: (2026)
by: Zhao, Zhongyin, et al.
Published: (2026)
On the Robustness of GUI Grounding Models Against Image Attacks
by: Zhao, Haoren, et al.
Published: (2025)
by: Zhao, Haoren, et al.
Published: (2025)
FreeVS: Generative View Synthesis on Free Driving Trajectory
by: Wang, Qitai, et al.
Published: (2024)
by: Wang, Qitai, et al.
Published: (2024)
Explorer: Robust Collection of Interactable GUI Elements
by: Chaimalas, Iason, et al.
Published: (2025)
by: Chaimalas, Iason, et al.
Published: (2025)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
by: Wang, Wenkai, et al.
Published: (2026)
by: Wang, Wenkai, et al.
Published: (2026)
ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding
by: Hsieh, ZongHan, et al.
Published: (2025)
by: Hsieh, ZongHan, et al.
Published: (2025)
ClickVOS: Click Video Object Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026)
by: Pei, Siqi, et al.
Published: (2026)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
by: Li, Junlong, et al.
Published: (2026)
by: Li, Junlong, et al.
Published: (2026)
DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model
by: Wang, Yuqi, et al.
Published: (2024)
by: Wang, Yuqi, et al.
Published: (2024)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
MVP: Multiple View Prediction Improves GUI Grounding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
CityGo: Lightweight Urban Modeling and Rendering with Proxy Buildings and Residual Gaussians
by: Liu, Weihang, et al.
Published: (2025)
by: Liu, Weihang, et al.
Published: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
by: Li, Weiming, et al.
Published: (2025)
by: Li, Weiming, et al.
Published: (2025)
ClickDiff: Click to Induce Semantic Contact Map for Controllable Grasp Generation with Diffusion Models
by: Li, Peiming, et al.
Published: (2024)
by: Li, Peiming, et al.
Published: (2024)
What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs
by: Lin, Jiaping, et al.
Published: (2026)
by: Lin, Jiaping, et al.
Published: (2026)
Fully Sparse Fusion for 3D Object Detection
by: Li, Yingyan, et al.
Published: (2023)
by: Li, Yingyan, et al.
Published: (2023)
Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
by: Zhao, Haoren, et al.
Published: (2026)
by: Zhao, Haoren, et al.
Published: (2026)
ClickRemoval: An Interactive Open-Source Tool for Object Removal in Diffusion Models
by: Zhang, Ledun, et al.
Published: (2026)
by: Zhang, Ledun, et al.
Published: (2026)
iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
ClickSeg3D: Few-Click Interactive Segmentation via Semantic Embeddings
by: Kang, Xueyang, et al.
Published: (2026)
by: Kang, Xueyang, et al.
Published: (2026)
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
by: Zhang, Shaojie, et al.
Published: (2025)
by: Zhang, Shaojie, et al.
Published: (2025)
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
Similar Items
-
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
by: Li, Hongxin, et al.
Published: (2025) -
UIPro: Unleashing Superior Interaction Capability For GUI Agents
by: Li, Hongxin, et al.
Published: (2025) -
AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark
by: Li, Hongxin, et al.
Published: (2026) -
DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
by: Chen, Yuntao, et al.
Published: (2024) -
Enhancing End-to-End Autonomous Driving with Latent World Model
by: Li, Yingyan, et al.
Published: (2024)