Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yan, Wu, Daiqing, Shen, Huawen, Ma, Can, Zhou, Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917476432871424
author Zhang, Yan
Wu, Daiqing
Shen, Huawen
Ma, Can
Zhou, Yu
author_facet Zhang, Yan
Wu, Daiqing
Shen, Huawen
Ma, Can
Zhou, Yu
contents Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capability for autonomous GUI agents. Recent reinforcement learning methods (e.g., GRPO) have achieved strong performance, but they rely on expensive multiple rollouts and suffer from sparse signals on hard samples. These limitations make on-policy self-distillation (OPSD), which provides dense token-level supervision from a single rollout, a promising alternative. However, its applicability to GUI grounding remains unexplored. In this paper, we present GUI-SD, the first OPSD framework tailored for GUI grounding. First, it constructs a visually enriched privileged context for the teacher using a target bounding box and a Gaussian soft mask, providing informative guidance without leaking exact coordinates. Second, it employs entropy-guided distillation, which adaptively weights tokens based on digit significance and teacher confidence, concentrating optimization on the most impactful and reliable positions. Extensive experiments on six representative GUI grounding benchmarks show that GUI-SD consistently outperforms GRPO-based methods and naive OPSD in both accuracy and training efficiency. Code and training data are available at https://zhangyan-ucas.github.io/GUI-SD/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00642
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
Zhang, Yan
Wu, Daiqing
Shen, Huawen
Ma, Can
Zhou, Yu
Artificial Intelligence
Computer Vision and Pattern Recognition
Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capability for autonomous GUI agents. Recent reinforcement learning methods (e.g., GRPO) have achieved strong performance, but they rely on expensive multiple rollouts and suffer from sparse signals on hard samples. These limitations make on-policy self-distillation (OPSD), which provides dense token-level supervision from a single rollout, a promising alternative. However, its applicability to GUI grounding remains unexplored. In this paper, we present GUI-SD, the first OPSD framework tailored for GUI grounding. First, it constructs a visually enriched privileged context for the teacher using a target bounding box and a Gaussian soft mask, providing informative guidance without leaking exact coordinates. Second, it employs entropy-guided distillation, which adaptively weights tokens based on digit significance and teacher confidence, concentrating optimization on the most impactful and reliable positions. Extensive experiments on six representative GUI grounding benchmarks show that GUI-SD consistently outperforms GRPO-based methods and naive OPSD in both accuracy and training efficiency. Code and training data are available at https://zhangyan-ucas.github.io/GUI-SD/.
title Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.00642